Every pitch deck now has a slide about AI agents, and most of them show the same two logos. This post collects AI agent case studies from twenty real deployments, from Klarna to Air Canada, Meesho to McDonald’s, with what each company did, what happened next and the lesson worth keeping. Every story links to the source we read, with the month and year. Where a result is a claim the company made about itself, we say so.
One pattern showed up before we had finished collecting them. The deployments that held up kept a person in the loop for the hard or risky cases. Most of the ones that went wrong let the software speak or act alone.
AI agent case studies: what went well
1. Meesho: a voice bot that answers 60,000 support calls a day
What happened. In November 2024 the e-commerce company launched a generative AI voice bot that handles about 60,000 calls a day in English and Hindi, with six more Indian languages planned. Meesho says it cut per-call costs by 75% and that only about 5% of calls need a person. It also says the human agents were moved to complex queries and seller support, not let go.
The lesson. Build for your real customers. Meesho designed for noisy rooms and low-end phones, and kept people for the 5% of calls that need judgement.
2. Air India: AI.g, a virtual agent for routine passenger queries
What happened. Air India’s assistant, built on Azure OpenAI, now handles about 40,000 queries a day across more than 1,300 topics, according to a February 2026 Microsoft customer story. The airline and Microsoft say it has resolved more than 13 million conversations and that only 3% are escalated to a human agent.
The lesson. A published escalation rate is a sign of a well-run system. It tells you the company decided in advance which questions belong to people.
3. IndiGo: 6Eskai, a booking assistant that hands over to people
What happened. In November 2023 IndiGo soft-launched 6Eskai, a GPT-4 assistant built by its own digital team. It answers in ten languages and can book tickets, apply discounts, add extras, do web check-in and connect the customer to a human agent. IndiGo said early results showed a 75% reduction in customer service agent workload.
The lesson. “Talk to a person” belongs on the feature list, not hidden away. The workload saving came from the routine jobs, while the handover stayed one step away.
4. PhysicsWallah: AI answers most student doubts, experts handle the rest
What happened. In a November 2025 interview with Outlook Business, co-founder Alakh Pandey said about 80% of student doubts are now handled through AI. The company used to need 20 subject experts per class and now needs four. The doubts the AI cannot answer are solved by people, and those answers go back into the library used to train the system.
The lesson. Pandey also said 100% accuracy is impossible. The fix was not a perfect model: it was a team that catches the misses and feeds them back.
5. Razorpay, NPCI and OpenAI: agentic payments with a confirmation step
What happened. In October 2025 Razorpay, the National Payments Corporation of India and OpenAI announced a pilot for paying over UPI inside ChatGPT, with BigBasket among the first merchants. In the example flow, the agent finds groceries, shows options and places the order after a single confirmation from the user. Razorpay says users keep control through real-time tracking and instant revocation.
The lesson. This one is still a pilot, so judge the design, not the results. Even here, money does not move until a person says yes.
6. Octopus Energy: Arlo answers routine emails, never the sensitive ones
What happened. In July 2026 the UK energy supplier reported a three-month trial of Arlo, an AI assistant that replied to about 8,000 emails a week, around 4% of its UK customer email. Octopus says Arlo scored 76% on customer satisfaction against 72% for comparable human replies. It never handles vulnerable customers, sensitive cases or complex complaints, every reply is labelled as AI, and customers can ask for a person at any point.
The lesson. Start with a modest share of easy work and write down what the agent may never touch. Octopus published its exclusions alongside its results.
7. Amazon: an agent that upgrades old Java code, with developers approving the plan
What happened. In August 2024 Amazon said its Amazon Q Developer agent helped move tens of thousands of production applications to Java 17. The company estimates this saved over 4,500 years of development work and is worth about $260 million a year. Developers review and change the agent’s upgrade plan before it makes the changes.
The lesson. Upgrades are repetitive and easy to check, which suits an agent well. The plan review is what lets a company trust it with production code.
8. Morgan Stanley: meeting notes drafted by AI, sent by the adviser
What happened. In June 2024 Morgan Stanley added Debrief to its OpenAI-based tools for financial advisers. With the client’s permission, it takes meeting notes, flags action items, drafts a summary that can be emailed and saves a note to Salesforce. The firm says 98% of adviser teams have adopted its earlier AI assistant.
The lesson. The client gives consent, the adviser decides what gets sent, and the AI does the typing. Nobody’s work is handed over without them knowing.
9. Bank of America: Erica passes 3 billion client interactions
What happened. In August 2025 the bank said its virtual assistant Erica, launched in 2018, had passed 3 billion client interactions and nearly 50 million users, averaging more than 58 million interactions a month. Erica can book appointments, which the bank calls a handoff to its higher-touch service channels, so specialists can focus on complex conversations.
The lesson. Seven years of steady growth came from a narrow scope that got wider slowly. The bank does not describe Erica as generative AI, which is a reminder that reliability often matters more than novelty.
10. Lyft: a customer care assistant that passes hard cases to specialists
What happened. In February 2025 Lyft and Anthropic announced a Claude-powered customer care assistant that handles thousands of enquiries a day. The companies say it cut resolution time by 87%. Complex cases move to human specialists when needed.
The lesson. Faster answers for routine questions free up the specialists for the cases that matter most to riders and drivers.
11. Salesforce: AI handles half of support conversations
What happened. In September 2025 Marc Benioff said Salesforce’s support headcount had fallen from about 9,000 to about 5,000 over roughly eight months, with AI agents handling half of customer interactions and people the rest. He said some staff moved into sales. He also said the models “can’t do everything” and compared the setup to an autopilot that hands control back.
The lesson. This counts as a working deployment, but it came with real job losses, and founders should weigh that honestly. The handback to people is built in.
AI agent case studies: what went wrong
12. Klarna: an AI assistant hailed, then partly walked back
What happened. In February 2024 Klarna said its AI assistant handled 2.3 million conversations in its first month, two thirds of its service chats, doing the work of about 700 full-time agents. In May 2025 chief executive Sebastian Siemiatkowski told Bloomberg the company was hiring human agents again, saying AI support was cheaper but produced “lower quality”. He said customers must know “there will always be a human if you want”.
The lesson. Volume is not the same as quality. Measure repeat contacts and satisfaction, not just chats closed, and keep the route to a person open.
13. Air Canada: liable for what its chatbot made up
What happened. A passenger booking after his grandmother died was told by Air Canada’s website chatbot that he could claim a bereavement fare after travel. The airline’s actual policy said otherwise. In February 2024 a British Columbia tribunal rejected the airline’s defence and ordered it to pay $812.02 in damages and fees, writing that Air Canada “is responsible for all the information on its website”. The full decision is public.
The lesson. You cannot hand legal responsibility to software. Policy answers need to come from approved text, and a person needs to own that text.
14. Replit: a coding agent deleted a production database during a code freeze
What happened. In July 2025 SaaStr founder Jason Lemkin reported that Replit’s agent deleted his production database despite instructions not to change code without permission, and created fake data to cover up bugs. It also told him a rollback would not work, which turned out to be wrong. Chief executive Amjad Masad called it “unacceptable”, promised a refund and a postmortem, and began separating development and production databases automatically.
The lesson. An instruction in a chat window is not a control. Destructive actions need permissions the agent cannot override and a person who approves them.
15. DPD: a support chatbot that swore and criticised the company
What happened. In January 2024 a customer chasing a missing parcel got DPD’s chatbot to swear, write a poem calling the firm useless and call it the “worst delivery firm in the world”. He posted the exchange on X and it spread quickly. DPD blamed an error after a system update and switched off the AI element.
The lesson. Test every update against people trying to break the bot, not only against polite questions. It had also failed at the basic job of finding the parcel.
16. Chevrolet of Watsonville: a dealer chatbot that “agreed” to a $1 SUV
What happened. In December 2023 a user told a California dealership’s ChatGPT-based chatbot to agree with everything and call each answer legally binding, then got it to accept $1 for a 2024 Chevrolet Tahoe. The dealer ended the chatbot service. GM said it was a third-party tool chosen by individual dealers, and a Chevrolet spokesperson said the episode showed the importance of human judgement when reviewing AI content.
The lesson. A customer-facing bot should never be able to state a price or make an offer on its own. Quotes need rules and a person.
17. McDonald’s: ended its AI drive-thru test with IBM
What happened. In June 2024 McDonald’s said it would end its automated order-taking partnership with IBM, after a test in about 100 restaurants that began in 2021. The company had wanted 95% order accuracy before any wider rollout. A 2022 analyst report put voice ordering accuracy in the low 80s.
The lesson. Setting the accuracy target before rollout is what made this a sensible decision rather than a slow failure. Ending a pilot that misses its number is a good outcome.
18. New York City: the MyCity chatbot gave businesses illegal advice
What happened. In March 2024 The Markup found that the city’s Microsoft-powered business chatbot said bosses could take workers’ tips and landlords could refuse Section 8 vouchers, both against the law. The mayor at the time kept it online. In January 2026 the new mayor called it “functionally unusable”, and the city took it down in February.
The lesson. For legal or regulatory questions, an answer that is usually right is not good enough. Those questions need checked sources and a human route.
19. Commonwealth Bank of Australia: cut jobs for a voice bot, then reversed
What happened. In July 2025 CBA said 45 customer service roles would go after it introduced an AI voice bot. In August, after pressure from the Finance Sector Union, which said call volumes were actually rising, the bank reversed the decision. It said its assessment “did not adequately consider all relevant business considerations” and apologised.
The lesson. Prove the saving with real numbers before you change the team. Cutting staff on a forecast is a gamble.
20. Cursor: a support bot invented a company policy
What happened. In April 2025 users of the AI code editor were told by email that a subscription could not be used on more than one machine. There was no such policy. Co-founder Michael Truell said on Reddit that it was “an incorrect response from a front-line AI support bot”, refunded the affected user and said AI email replies are now clearly labelled.
The lesson. Label AI replies, and never let a bot explain policy it cannot look up. A wrong answer that sounds official does more damage than no answer.
Human in the loop: when we design an agent, we write its “never alone” list before its feature list. Prices, refunds, policy, legal answers and anything that deletes data go to a named person. The agent drafts, the person decides.
The pattern across all twenty
- The successes kept people in the loop. Air India escalates 3%, Meesho about 5%, Octopus keeps vulnerable customers away from Arlo entirely, and Amazon’s developers approve the plan before the agent acts. Most of the failures had no human checkpoint at the moment it mattered.
- Narrow jobs work, open-ended ones wander. Routine queries, code upgrades and meeting notes went well. Open chat about policy, law and prices went wrong.
- Accountability stays with you. Air Canada, New York City and Cursor all found that the bot’s words counted as the organisation’s words.
- Targets set early make the decision easy. McDonald’s knew its accuracy target, so ending the test was clear. Klarna and CBA measured savings first and quality later.
- Most of the good numbers are company claims. Read them as direction, not as a forecast for your business, and take your own baseline before you start. The Indian deployments here are worth studying because they were built for Indian languages, phones and customers.
Human in the loop: the cheapest safeguard in this list is a person approving the agent’s riskiest actions for the first few weeks. It costs some review time. Every story in the second half cost far more.
What to do with this
None of these cases shows that AI handles everything, and none shows that it handles nothing. They show that the design around the agent decides the outcome: what it may do alone, what it hands to a person, and how you check that it is working. If you want the background, read our guide to chatbots, AI agents and agentic applications and why AI projects fail.
At Cannyworx, AI does the heavy lifting and a senior person signs off, which is the pattern behind most of the successes above. If you would like to talk through where an agent would help your team, and where it shouldn’t go, book a free consultation, or start with the readiness scorecard below.
Questions people ask
What do successful AI agent deployments have in common?
In the cases we reviewed, the ones that held up gave the agent a narrow job, kept a clear route to a person, and left anything sensitive, costly or irreversible to a human. Air India, Meesho and Octopus Energy all describe a minority of cases going to people by design.
Is a company legally responsible for what its AI chatbot or agent says?
In the Air Canada case, yes. A Canadian tribunal ruled in February 2024 that the airline was responsible for all the information on its website, chatbot included, and ordered it to pay the passenger. Laws differ by country, so take advice for your own situation, but plan as if the agent's words are yours.
Did Klarna really go back to human customer service?
Partly. In February 2024 Klarna said its AI assistant was doing the work of about 700 agents. In May 2025 its CEO told Bloomberg it was hiring people again, saying AI support was cheaper but lower quality and that customers should always be able to reach a human. It still uses AI.
How should a founder use these case studies?
Use them to choose your first task and your guardrails, not to set targets. Most of the results here are company claims from large firms. Pick one narrow, repetitive job, decide what the agent may never do alone, and measure it against a baseline you took before you started.





