The a16z Podcast: Decagon's Playbook for Building Enterprise AI Applications
The a16z Podcast: Decagon’s Playbook for Building Enterprise AI Applications
Sarah Wang and Kimberly Tan with Jesse Zhang and Ashwin Sreenivas. Duration: 1h 20m
Timestamps 1:34 Transitioning to open-source models for enterprise AI 14:44 Redefining the application layer and forward-deployed engineering 29:01 Automating agent development with autonomous systems 34:22 Enterprise sales strategy and the glass box approach 44:22 Scaling culture and international expansion 1:01:54 The future of enterprise software and AI concierges 1:14:17 AI’s impact on jobs and the future of work
Decagon reached a $1.5B valuation in about a year out of stealth by putting AI agents in front of real customer support desks at some of the world’s biggest companies. Jesse Zhang, its CEO, and Ashwin Sreenivas, its president, walk through the operational reality of running agents in production. The headline: they moved most of their inference to open-source models, and it wasn’t a cost cut, it was a quality decision.
Frontier models aren’t always the right answer in production. Decagon moved most of its inference to open-source models because latency, throughput, and control matter more than raw benchmark performance when a customer is waiting on the line. A smarter model that’s too slow is worse than a faster one that’s good enough.
The application layer is the moat, not the model. The hard part of enterprise AI isn’t calling an API, it’s encoding business logic: which workflows an agent may run, what a refund looks like for this company, when to escalate. That’s where Decagon builds its software.
Agents are becoming the front door of the enterprise. Decagon’s agents handle end-to-end customer interactions across chat, phone, SMS, and email. The customer’s first contact is increasingly an agent that can actually resolve the issue, not a chatbot that bounces you to a human.
Instructions beat decision trees. Earlier chatbots broke the moment a question left the script. Decagon’s agents follow instructions the way a trained person would, so a request one degree off the happy path still gets handled instead of failing.
Fine-tuning small models wins on latency. Decagon runs a network of specialized models, one for intent, one for workflow execution, one for hallucination detection, one for escalation. A fine-tuned small model gives you speed and reliability a single large model can’t guarantee, especially for voice.
Evaluation is the product. Every component gets continuously measured on accuracy, latency, resolution rate, and customer satisfaction. That eval loop is what lets Decagon swap models and fine-tune without regressing quality.
Forward-deployed engineering is a bridge, not a lifestyle. Early adopters need engineers embedded on-site to make the system work. The goal is to productize those custom workflows into software so the company stops needing to hold the customer’s hand.
Selling to the enterprise means showing your work. Big companies won’t hand over customer relationships to a black box. The “glass box” approach, making agent behavior and escalation visible to buyers, is how Decagon closes deals.
AI expands what work covers more than it deletes it. Jesse and Ashwin’s take on jobs: AI takes the mundane, repetitive tasks and frees people for higher-value, revenue-generating work. It’s part of the pitch, and it’s how they frame the future of support teams.
Application companies thrive alongside the model labs. The labs keep raising the ceiling, but someone has to build integrations, observability, and industry-specific workflows on top. Decagon’s bet is that the application layer gets more valuable as models get better.
Related TMFNK Content
- Lenny’s Podcast: How 80,000 Companies Build with AI: Asha Sharma (Microsoft) The other side of the enterprise AI table: how Microsoft sees companies actually adopting AI tools in production.
- Cheeky Pint: Cognition CEO Scott Wu on AI Agents Another CEO shipping autonomous agents into real workflows, from the Devin side of the agent stack.
- Latent Space: Inside the Model Factory with Eiso Kant (Poolside) Where the open-source models Decagon runs actually come from, straight from the people training them.
- 20VC: Groq Founder Jonathan Ross on Why Latency Wins The hardware argument for the low-latency inference Decagon’s agents depend on.
Crepi il lupo! 🐺