Why AI Infrastructure Must Evolve for Agent Experience: Lessons from Modal's Agent Cloud
TL;DR
- Traditional cloud infrastructure wasn't designed for AI agents: Agents need sub-second cold starts, persistent state management, and streaming capabilities that serverless and container platforms struggle to provide at scale.
- The "agent experience" requires rethinking the entire stack: From scheduling and orchestration to observability, infrastructure must support long-running, stateful, and unpredictable workloads rather than short-lived request-response patterns.
- Modal's agent cloud architecture prioritizes developer experience: By combining fast cold starts (200-400ms), native streaming support, and simplified state management, Modal demonstrates how infrastructure can remove friction from agent development.
- Product builders should demand more from their infrastructure: The gap between what agents need and what current platforms offer represents both a constraint on innovation and an opportunity for those who architect around agent-first principles.
We're at an inflection point in AI product development. The industry has spent the past two years obsessing over model capabilities—context windows, reasoning benchmarks, multimodal understanding. But there's a growing realization that infrastructure is now the bottleneck.
Akshat Bubna, CTO of Modal, articulated this shift in a recent conversation on the Latent Space podcast where he outlined why Modal rebuilt their platform specifically for agent workloads. His insights matter because Modal processes millions of AI workloads daily, giving them a front-row seat to how agent architectures actually behave in production.
The core thesis? The infrastructure we built for web applications and even for batch ML workloads fundamentally mismatches what AI agents need. And if you're building agent-centric products, understanding this mismatch isn't optional—it's existential.
The Hidden Constraints of Traditional Cloud Infrastructure
Most product builders inherit assumptions from decades of web infrastructure evolution. We've been trained to think in terms of stateless functions, request-response cycles, and horizontal scaling. These patterns work brilliantly for web apps. They fail spectacularly for agents.
Consider the typical agent interaction: A user asks a question. The agent needs to:
- Retrieve relevant context from vector databases
- Make multiple LLM calls with reasoning steps
- Execute tool calls to external APIs
- Stream partial results back to the user
- Maintain conversation state across turns
- Potentially spawn sub-agents for parallel work
This workflow violates nearly every assumption of traditional serverless platforms. It's not stateless (agents need memory). It's not short-lived (complex reasoning takes seconds or minutes). It's not predictable (tool execution depends on external systems). And it absolutely requires streaming (users won't wait for complete responses).
Bubna's team discovered this through direct observation. When developers tried to build sophisticated agents on Modal's original platform, they hit walls. Cold start times that seemed acceptable for batch jobs became deal-breakers for conversational experiences. The lack of native streaming support forced awkward workarounds. State management required external databases, adding latency and complexity.
What "Agent Experience" Actually Means
The term "agent experience" might sound like marketing speak, but it represents a genuine paradigm shift. It's not about the end-user experience (though that matters). It's about the developer experience of building systems where AI agents are first-class citizens.
Traditional infrastructure treats compute as the primitive. You request a container or a function invocation, run your code, and return a result. Agent infrastructure needs to treat the agent's lifecycle as the primitive.
This means infrastructure must natively understand:
Long-running, stateful sessions: Agents aren't fire-and-forget functions. They maintain context, learn from interactions, and evolve over multiple turns. Infrastructure needs to keep agents "warm" without burning resources on idle time.
Streaming as default, not an afterthought: Every agent interaction should stream results. Not because it's trendy, but because latency perception matters enormously in conversational interfaces. Infrastructure that makes streaming hard is infrastructure that makes good agent UX hard.
Unpredictable resource patterns: Agents might need 100ms of compute for a simple query or 30 seconds for complex reasoning with multiple tool calls. Traditional auto-scaling can't handle this volatility well—it's either over-provisioned (expensive) or under-provisioned (slow).
Observability that matches agent cognition: Logs and metrics designed for web requests don't map to agent behavior. You need to trace reasoning chains, understand tool call patterns, and debug multi-step workflows. Standard APM tools fall short.
Modal's response was to rebuild their scheduling layer, introduce native streaming primitives, and optimize cold starts to the point where they're barely perceptible (200-400ms even for GPU workloads). These aren't incremental improvements—they're architectural decisions that prioritize agent workloads over traditional cloud patterns.
My Take: Infrastructure Shapes What We Can Build
I think we dramatically underestimate how much infrastructure constraints shape the products we build. This isn't just about performance—it's about what ideas feel feasible.
When cold starts take 5-10 seconds, you don't build conversational agents. You build batch processors and convince yourself that's what users want. When streaming requires complex WebSocket management and state synchronization, you ship traditional request-response interfaces and tell yourself streaming is a nice-to-have.
Infrastructure doesn't just enable products; it defines the solution space we explore.
I've watched product teams abandon sophisticated agent architectures not because the AI wasn't capable, but because the infrastructure made iteration painful. Every debugging cycle took minutes. Every deployment risked cold start regressions. Every new feature required wrestling with state management across distributed systems.
The teams that succeed with agents today are often those who either:
- Have the resources to build custom infrastructure (large tech companies)
- Constrain their agent designs to fit existing infrastructure (most startups)
- Find platforms that actually understand agent workloads (still rare)
This creates a selection pressure toward simpler, less capable agents—not because that's what the market needs, but because that's what the infrastructure permits. Modal's bet is that by removing these constraints, they'll enable a new class of agent applications that currently seem impractical.
I think they're right. The most interesting agent products I've seen in the past six months all share a common trait: they're built by teams who either control their infrastructure stack or have found platforms that don't fight them.
The Technical Shifts That Matter for Product Builders
If you're building agent-centric products, here are the infrastructure capabilities you should be demanding (or building):
Fast Cold Starts Are Non-Negotiable
Bubna's team achieved 200-400ms cold starts even for GPU workloads through aggressive optimization of container scheduling and image caching. This matters because it changes the economics of keeping agents warm.
With 10-second cold starts, you're forced to keep instances running 24/7 or accept terrible UX. With sub-second cold starts, you can scale to zero aggressively and only pay for actual usage. For product builders, this means you can experiment with more sophisticated agent architectures without worrying about idle costs.
Native Streaming Throughout the Stack
Streaming shouldn't be something you bolt on. Modal introduced streaming as a first-class primitive, meaning you can stream results from any function without managing WebSocket connections, handling backpressure, or coordinating state.
For product builders, this means you can default to streaming UX patterns. Every agent response can show progressive results. Every long-running task can provide status updates. This transforms perceived latency and makes complex agent workflows feel responsive.
Simplified State Management
Agents need memory, but managing distributed state is traditionally complex. The right infrastructure abstracts this without sacrificing control.
Modal's approach includes built-in primitives for persistent state that survive across function invocations. You're not managing Redis connections or worrying about state synchronization—you declare what state you need, and the platform handles the rest.
For product builders, this means you can focus on agent logic rather than infrastructure plumbing. Your agents can maintain conversation history, learn from interactions, and coordinate across multiple turns without you becoming a distributed systems expert.
Observability That Matches Agent Workflows
Debugging agents is fundamentally different from debugging web apps. You need to understand reasoning chains, trace tool calls, and identify where in a multi-step workflow things went wrong.
The best agent infrastructure provides observability that maps to how agents actually work: structured logs that capture reasoning steps, trace visualization that shows tool call patterns, and metrics that reflect agent-specific concerns (reasoning time, tool success rates, context utilization).
The Broader Implications for Agent Product Strategy
The infrastructure conversation isn't just technical—it has strategic implications for how you build and position agent products.
Time-to-value matters more than ever: With fast cold starts and simplified state management, you can ship agent prototypes in hours rather than weeks. This changes how you validate ideas and iterate with users.
Cost structures shift dramatically: Agent workloads have different cost profiles than traditional apps. They're bursty, unpredictable, and often GPU-intensive. Infrastructure that scales to zero and charges only for actual usage makes agent economics viable at smaller scales.
Competitive moats may live in execution speed: When everyone has access to similar models, the winners will be those who can iterate fastest on agent architectures. Infrastructure that reduces friction in the development loop becomes a competitive advantage.
The multi-agent future requires better orchestration: As agents become more sophisticated, you'll increasingly build systems of cooperating agents rather than monolithic assistants. Infrastructure needs to make agent-to-agent communication, coordination, and resource sharing natural rather than complex.
What This Means for Your Next Agent Product
If you're starting a new agent-centric product or refactoring an existing one, here's how to think about infrastructure:
Audit your current constraints: What agent architectures have you ruled out because they seemed too complex to implement? Which features would you build if cold starts weren't an issue? Where does state management create friction?
Prototype with agent-first platforms: Before committing to custom infrastructure, test whether platforms like Modal, Inngest, or agent-specific offerings remove your constraints. The ROI of not building infrastructure is enormous if the platform matches your needs.
Design for streaming from day one: Don't treat streaming as a v2 feature. Build your agent UX around progressive results and status updates. Users will forgive latency if they see progress.
Instrument for agent-specific metrics: Track reasoning time separately from tool execution time. Measure tool success rates and context utilization. These metrics matter more than traditional web metrics for agent products.
Plan for multi-agent architectures: Even if you're starting with a single agent, design your infrastructure to support agent coordination. The most sophisticated agent products will be ecosystems of specialized agents, not monoliths.
The Infrastructure Gap Is Closing
The good news is that the infrastructure gap for agent workloads is finally getting attention. Modal's agent cloud is one example, but we're seeing movement across the industry: vector databases adding agent-specific features, observability platforms building agent tracing, and orchestration tools designed specifically for multi-agent systems.
For product builders, this means now is actually the time to build ambitious agent products. The infrastructure constraints that made sophisticated agent architectures impractical six months ago are rapidly disappearing.
The teams that win in the agent era won't necessarily be those with the best models or the most data. They'll be the teams that understand how to architect agent systems that feel fast, reliable, and natural to users—and that requires infrastructure that was purpose-built for how agents actually work.
Bubna's insights from building Modal's agent cloud provide a roadmap. The question is whether you'll wait for infrastructure to evolve, or actively seek out (or build) the platforms that let you ship the agent products you actually want to create.
The agent experience revolution isn't just about what AI can do. It's about building the infrastructure that lets us discover what's possible when we stop fighting our tools and start building with them.
Frequently Asked Questions
What makes AI agent workloads different from traditional web application workloads?
AI agent workloads are fundamentally stateful, long-running, and unpredictable compared to traditional stateless web requests. Agents need to maintain conversation context across multiple turns, execute complex reasoning chains that can take seconds or minutes, make unpredictable tool calls to external systems, and stream partial results continuously. Traditional serverless and container platforms were optimized for short-lived, stateless request-response patterns and struggle with these requirements, leading to poor user experiences and difficult development workflows.
Why are fast cold starts so important for AI agent infrastructure?
Fast cold starts (under 500ms) change the economics and user experience of agent applications. With slow cold starts (5-10 seconds), you're forced to keep instances running continuously to avoid terrible UX, which is expensive and wasteful. Sub-second cold starts allow you to scale to zero and only pay for actual usage while maintaining responsive user experiences. This makes it economically viable to build sophisticated agent architectures without worrying about idle infrastructure costs, and enables product teams to experiment more freely with complex agent designs.
How should product teams evaluate whether they need agent-specific infrastructure?
Start by auditing your current constraints: identify agent architectures you've ruled out due to implementation complexity, features you'd build if cold starts weren't an issue, and places where state management creates friction. If you're building conversational agents, multi-step reasoning systems, or applications where agents coordinate with each other, agent-specific infrastructure will likely provide significant advantages. The key indicators are: struggling with streaming implementations, spending significant time on state management, or finding that infrastructure concerns are limiting your agent design choices rather than enabling them.
What observability capabilities matter most for debugging agent systems?
Agent systems require observability that maps to how agents actually reason and act, not just traditional logs and metrics. Essential capabilities include: structured logging that captures reasoning steps and decision points, trace visualization that shows tool call patterns and dependencies, metrics that reflect agent-specific concerns like reasoning time versus tool execution time, and the ability to understand multi-step workflows across agent interactions. Standard application performance monitoring tools designed for web requests typically fall short because they don't provide visibility into the cognitive patterns and coordination behaviors that define agent performance.