An open letter to Doordash labs: the job Amazon could not automate away
What fourteen years and a million robots cost Amazon, and what that record predicts for DoorDash Air. From someone who helped lay the foundation.
Stanley,
There is a job at Amazon called Amnesty Floor Monitor.
Amnesty is what Amazon calls an item that ends up loose on the floor of a robotic field, blown there by the airflow of a few hundred drive units in motion or pushed out of a pod by an arm that failed its stow. The Amnesty Floor Monitor keeps the floors clear and resets units when needed. That is the whole role. A person, in a safety vest whose green light tells nearby drives to hold, walking a caged grid, picking up what the automation dropped.
I know that role well, because I was there when we invented it. In 2012 we called it Flow Assistant.
Amazon was a customer of Kiva before it became Kiva’s owner. I was at Amazon on the team that ran those initial deployments, from 2011, through the acquisition in March 2012, and into the build of the first fully autonomous fulfillment center. No one’s design docs included Flow Assistant. We found we needed it in the first weeks of live operation, because inventory came off pods and ended up on the floor. And a drive unit that meets a fallen item is a drive unit that stops, and a drive unit that stops in a dense grid is the first bead in a jam that propagates. So we put a person down on the floor. This was meant to be a temporary thing.
Fourteen years later, it has an official title, a line for headcount, and a published failure rate.
In a paper covering March 2025, Amazon’s robotic stow team reported that roughly 3.7% of stow attempts pushed an item onto the floor. Stow success was about 85% across more than 500,000 attempts. The companion pick robot succeeded on 91% of what it attempted but declined 19.4% of the requests it was handed. On the same floor, in the same month, human associates stowed 243 units an hour against the robot’s 224.
Those numbers are from the most advanced applied physical AI deployment on earth, thirteen years and a million units in. They refer to a system that is slower than a person, rejects one job in five and spews out a steady stream of debris that someone is paid to clear up.
That system is a victory. I’m not being sarcastic. It is the right model to study. But here is the finding I want to present to you, and it’s why I’m writing instead of watching:
The recovery role was not removed. Not more than fourteen years. Six generations of hardware. A thousandfold increase in fleet size. A company with functionally unlimited capital and the best applied robotics team in commerce. It was renamed, regularized, staffed and metered. It was never engineered out.
Your team will have someone who will tell you that your equivalent of the Flow Assistant is transitional, scaffolding that you retire when the stack matures. We thought that in 2012. And we were wrong. And we have been wrong for fourteen years.
What you actually bought on wednesday
First, the congratulations, and it is specific. DoorDash is the eighth drone operator to earn Part 135 air carrier certification in more than a decade of attempts. That’s an organizational achievement, not a hardware one. Someone wrote an operations manual, developed a maintenance program, brought in a director of operations and chief inspector, established a records system and passed a five-step FAA evaluation. Most engineering organizations do not produce that set of documents. That’s why there are eight names on the list.
Now for the uncomfortable part. A Part 135 certificate is not a permission slip. It’s a standing obligation whose maintenance cost scales with the number of different operational configurations you run, and not at all with the number of deliveries you make. Every new aircraft type, base and material procedure change nibbles at the scarcest resource you have now - the attention of the few people who can hold your safety case in their heads and defend it to a regulator.
You didn’t buy scale. You bought an obligation and moved your binding constraint to a place where you have less experience running it.
The arithmetic under your own diagnosis
On No Priors you said the software stack is no longer the primary barrier, that the real ones are operational integration and hardware manufacturing, and you gave the sharpest possible example: fleet boot scripts that work for ten robots and break at five hundred.
That’s not a code quality problem. p is the probability that a unit will not be ready to go on any given morning, for any reason. Stale certificate, DHCP lease, timeout in the merchant portal, firmware skew, charger that said full but wasn’t... At ten units and p = 2% an exception occurs every fifth morning and one person takes it with a phone. At five hundred units and the same p, ten exceptions come every morning, before the breakfast rush, across a metro. You can’t get that on a phone.
The per-unit reliability required for constant human attention scales with fleet size. The operator’s morning is still good. Five to ten hundred. P needs to be fifty times better. Nothing scripting gives fifty times.
Amazon ran the natural experiment. The fleet went from roughly 200,000 drive units in 2019 to over a million in July 2025 across 300-plus facilities. If robots substitute for the humans who tend them, support headcount per facility should have fallen.
It rose.
Amazon says its new generation of fulfillment centers need 30% more workers in reliability, maintenance and engineering jobs. For sourcing hygiene: 25% according to Tye Brady at TechCrunch, 30% according to Amazon’s written announcement. I cannot reconcile them, nor will I deny that there is a discrepancy. In either case the sign is the same, and the sign is the finding.
So Amazon didn’t get a p low enough to pull the human out of the recovery path. It did two things at once. It drove p down hard, and built a permanent human organisation around the residual.
Its leaders tell the truth about this. “If we had to have Vulcan do 100% of the stows and picks, it would never happen,” says Aaron Parness, who leads the science behind Vulcan, Amazon’s first robot with tactile sensing. In that framing, the 19.4% pick-rejection rate is no defect. It is a work of design. The system knows what it is to avoid.
The two amazon programs that are already yours
Scout is Dot. Prime Air is Air. Amazon ran both of your programs, in the order you are running them.
Scout launched in January 2019, a cooler-sized six-wheeled sidewalk robot, expanding to Irvine, Atlanta, and Franklin, Tennessee. Each unit was accompanied by a human Scout Ambassador. In October 2022 Amazon ended field tests and disbanded the team of roughly 400 people, citing aspects of the program that were not meeting customers’ needs.
See what didn’t work and what did. The robot did its job. For nearly four years it drove sidewalks in four climates. What didn’t work was the economics of a machine that required a person to walk beside it.
Dot isn’t a Scout. It runs roads and bike lanes at 20 mph instead of crawling a sidewalk, it carries six pizza boxes, and it is dispatched by a network that already knows what a good handoff looks like from billions of deliveries. These are the proper differences. But Scout is the precedent that a robot that can’t do the last ten feet without a companion is capability complete and network incomplete, and the gap between the two states can absorb four years and four hundred people before anyone calls it.
The question is not whether Dot can drive. It certainly can. What is your ratio of human interventions to completed deliveries, what is its trend across successive waves of deployment and below what number do the unit economics close? It is: If no one can answer the third part, then the program doesn’t have a definition of done yet.
Prime Air has been in operation since 2013. As of February 2026, it had made about 16,000 lifetime deliveries, nowhere near its stated target of 500 million a year by the end of this decade. “Between one and thirty thousand somewhere.
What is useful is the history between those numbers. There were at least eight test crashes in a twelve-month window in the MK27 era, one starting a brush fire. Mid-air crash during engine failure simulation. Two crashes in Pendleton December 2024. Two-month service break. Resuming after dust in dry Arizona air, a condition Texas never had, impacted altitude sensing fixes. In October 2025, two aircraft struck the same construction crane in Tolleson within minutes of each other. In November, one severed a residential internet cable in Waco thirteen days after launch. In February 2026, one hit an apartment building in Richardson.
The validation distribution did not have environmental or temporal domain shift. Most of these are one failure in different clothing: A regime of precipitation. A system of particles. Same-day change in the vertical obstacle geometry.
The crane Tolleson wants is the one to print out. There is no aircraft failure at all. Two aircraft encountered the same new problem minutes apart, with the second not knowing what the first had just learned. That is a breakdown of fleet knowledge. You have one vehicle, and it does not disappear by making the vehicle better. It does not exist.
Two asks come out of that decade. Propagate discovered obstacles fleet-wide, in minutes, not release cycles. And fund environmental domain coverage as a separate budget line from behavioral coverage, because engineers naturally enumerate the construction zone and the off-leash dog, while environmental regimes are what actually ground fleets and are boring enough that nobody volunteers to own them.
The argument for the position you already hold
You said the future is more Dashers, not less. I think more properly than you put it.
Weather degrades air capacity in an ungraceful manner. It turns it off, simultaneously, spatially. And in the worst possible way, the covariance. Rain grounds aircraft, rain reduces Dasher supply because fewer people want to drive in the rain, rain raises demand because fewer people want to leave the house. At the point where the air drops to zero, ground fallback is thinnest and demand highest.
The only mode of which availability is imperfectly correlated with all others is human supply. That is therefore the reserve asset that prices the risk in each autonomous mode. If anything, that argument only gets stronger as autonomy scales, and the proof is Amazon’s history: a million robots in and the human organization around them grew.
In practice: price reserve against joint distribution of mode availability, Dasher supply and demand. In code, make sure no order goes to a mode whose abort has no ground fallback inside the food-quality window. Weather is not an availability haircut. Don’t model it that way. A haircut is an independence assumption, and independence is the one thing weather is not.
One question for your leadership team
In the footage you put out on Wednesday there is a shot of the workshop. Benches, parts bins, cable spools, an airframe in pieces, a rotor spinning inside a test cage. I recognize that room. Ours looked like it in 2011, and I mean that as a compliment, because it is the correct room for this stage. The building with LABS painted on the roof is your sandbox.
We didn’t grow Amazon by adding robots. We scaled through proving a facility. Sandbox, template, replicate, with discipline living entirely in the freeze. A template that is changing all the time is a permanent sandbox with a bigger budget, and the replication economics never come. Thirteen years later, the Amazon unit of scaling is still the building, multiplied across some forty next generation sites by the end of 2027.
The sandbox is not the achievement. The specification that leaves it is, along with whether anyone is allowed to keep changing it afterward. Which is the question:
What’s your template unit? And have you frozen it?”
Conveniently there is a facility with walls and so for Amazon it is a facility. Your network doesn’t have one so the template unit is a design choice, not a physical fact, and the wrong choice is invisible for two years. A metro is the obvious answer and I think it is wrong, because metros are different in ways that you cannot control. Site archetypes make better candidates: the drive-through, the strip-mall back door, the rooftop, the apartment lobby, the suburban driveway. Those are repeated in metros. Freeze a handful, prove the handoff for each one. Then a new metro is a composition problem, not a discovery problem.
One more property, as it determines how carefully you choose the first ones. The sandbox is not permanent. The site where a system is first made to work accrues operating hours and institutional knowledge that no other site can buy, and it tends to stay among the strongest in the network long after everyone stops calling it the experiment. Pick your first archetype as if you’ll still be running it in 2036, because you probably will.
The one thing to do
Instrument the recovery layer now while it is a line item not a crisis and give it an owner senior enough to be inconvenient.
Two numbers, published weekly, at same seniority as delivery volume :
a) Human minutes per day of unit cause based breakdown. Your real cost of goods. And the real test of whether these assets are autonomous.
b) Interventions per finished delivery and its trend over deployment waves. If that number does not go down as you scale, you are still capability-limited, thinking you are network-limited.
Neither will be important until the quarter where it is the only thing people are talking about. And that is exactly why it needs a person and not a working group. DoorDash Labs is a startup inside an enterprise, which is the right structure, the same structure that produced everything Amazon has in robotics, but it has one characteristic failure mode: a startup inside an enterprise is measured on shipping, and recovery infrastructure never ships. No announcement, no demo, no release date. It does not drop arguments. It never gets into anybody’s objectives, so it never gets into anybody.
Claims you can grade me on
By the end of 2027, the fully loaded cost per successful delivery for DoorDash Air will match the equivalent three-to-five-mile Dasher delivery. The crossover point is regulatory, not technological. As of 10 July 2026 Part 108 was still at OIRA. Confidence 75 per cent.
In terms of Air throughput, your first three metros will be constrained by ground handling and pad service time first and foremost, not aircraft availability or autonomy capability. Confidence 80%
DoorDash Air is expected to have at least one voluntary or regulator-prompted operational pause of seven days or longer before the end of 2028. Not a critique. A baseline rate. Confidence 65%
Support and recovery personnel per active autonomous unit will not decrease monotonically through 2028. The direct transfer of the Amazon finding, and the one I most hope I am wrong about. Confidence 70%.
No US autonomous delivery operator will publicly disclose an autonomous fault resolution rate or a human-minutes-per-unit figure before the end of 2028. The operational layer gets named freely in interviews and disclosed almost nowhere, which is itself the finding. Confidence 85%.
The most famous technical letter ever sent to an American president concluded that such bombs might prove “too heavy for transportation by air.” It was wrong, and it looked wrong after the fact, because the author put the assumption on the page instead of burying it. That is the bar. Rate me.
What is at stake
The prize is not a drone program. It’s the stance your announcement quietly makes: the mode-agnostic clearing layer for physical local commerce. When the aircraft does, or the sidewalk robot does, or the driving model does, whoever can decide in under a hundred milliseconds what physical mode moves an object for the last three miles has something that doesn’t commoditize.
Inexpensive airframes. People building autonomy stacks will find them getting cheap faster than they expect. What’s scarce is the stuff that intelligence can’t make, that took ten years to build. The certificate. The safety case. The pad. The merchant relationship. The demand forecast. The trust of a city council. The technician who can turn an aircraft around in eleven minutes. The plane is the denominator. The numerator is the ground.
“You put it in a frame.” The plane is not the whole ball of wax. That is the airline. That’s the harder work, the less glamorous work, and the work your current company is better positioned to do than any specialist you now compete with. That asymmetry is real and it won’t be permanent. I would use it.
I may be wrong on some of this. You have never seen my internal numbers, and I am reasoning from the public record and lessons bought at another company, another decade, on a warehouse floor rather than in the air. If the number of interventions is already on a downward trend, I have written you a long letter about a problem you solved last quarter and I would rather hear that than be right .
But if any of it does hit, the boring one is the thing to do. Not the airplane.
Go find out what you are already calling your Flow Assistant. Not so you can hire more of them. So you know how much of the work still needs a person to walk over to the machine, and can start designing that walk out.
Again, Congratulations. The certificate is a very hard thing, honestly earned.



