
Modal Labs
AI Inference Infrastructure
Last verified August 15, 2026 · Updated daily
What Modal Labs is building
The workflow for deploying an AI model into production is deceptively hard. A data scientist or ML engineer finishes training a model on their local machine or a cloud notebook. They then hand it to a platform engineering team, who containerises it, writes Kubernetes configuration files, sets up autoscaling rules, configures a load balancer, establishes monitoring, manages GPU quota requests, and wires it into the company's existing infrastructure. That process takes days to weeks, requires specialised infrastructure knowledge that most ML practitioners don't have, and must be repeated in some form every time the model changes, the load profile shifts, or a new use case requires a different hardware configuration. The model itself might have taken an afternoon to train. The infrastructure work to run it reliably takes a sprint. Modal abstracts all of this. You write a Python function, add a @modal.function decorator, and Modal handles everything else: container construction, GPU provisioning, autoscaling from zero to hundreds of parallel workers, cold start optimisation, job scheduling for background tasks, secret management, and monitoring. The billing model matches the architecture: you pay per second of actual compute used, with no idle costs when nothing is running. For a team running inference workloads that spike unpredictably, a common pattern in production AI applications, the cost difference versus a reserved GPU cluster can be 60–70%. Modal builds and caches container images that start in under a second, compared to the 30–90 second cold starts that make naive containerised GPU workloads unusable for latency–sensitive applications. Their custom network file system allows large model weights to be loaded into GPU memory faster than any cloud provider's native storage solution. Their sandboxing architecture lets untrusted code run safely at scale, which enables use cases like code execution environments and customer–facing AI agents that cloud providers' general–purpose compute cannot safely support. Erik's insight: the biggest bottleneck in ML deployment isn't model quality – it's the gap between 'it works on my laptop' and 'it runs reliably at scale.' Modal closes that gap with a developer–first interface.
Why this matters
AI Training, the phase that dominated AI infrastructure spending from 2020 to 2024, involving massive GPU clusters running for weeks to produce a single model, is becoming a smaller fraction of total AI compute spend. Inference, or running trained models to generate outputs for actual users and applications, is becoming the dominant cost centre, and it is growing faster than training ever did because every new AI application generates continuous, ongoing inference demand from the moment it ships. For context, OpenAI processes over 10 billion words of inference output per day. Every enterprise that deploys an internal AI assistant, every consumer app that integrates a language model, every production system that runs AI–powered decisions generates inference compute demand continuously, around the clock, at volumes that compound as adoption grows. The current solutions are inadequate in ways that are becoming increasingly expensive. Hyperscaler GPU clouds – AWS, Google Cloud, Azure – were built for general–purpose compute workloads and adapted for AI inference. They are expensive, operationally complex, and optimised for reserved capacity rather than the spiky, unpredictable demand profiles that characterise real AI applications. A company running a consumer AI product cannot predict whether it will need 10 or 1,000 GPU instances at 3pm on a Tuesday – and paying for reserved capacity to handle the peak means paying for idle capacity during the troughs. The economics are structurally inefficient for the actual usage pattern of production AI. Modal's serverless abstraction converts infrastructure time into model time, a trade that, at the margin, is worth more than the raw compute cost difference, because ML engineer time is priced at $200–300K annually and infrastructure debugging compounds into weeks per quarter. The competitive dynamics of the inference infrastructure market are moving fast. Baseten, Modal's closest comparable, raised at a $5B valuation in late 2025. Together AI, Replicate, and Fireworks AI are all pursuing adjacent positions. The hyperscalers are shipping dedicated inference products. But the developer–first positioning – the bet that the primary buyer of inference infrastructure is the individual ML engineer or small team, not the enterprise procurement committee – is Modal's specific angle, and it reflects Erik Bernhardsson's background as a practitioner who built tools for other practitioners rather than a founder who came up through enterprise sales. The $50M ARR at Modal's current scale suggests that positioning is working. The question the General Catalyst raise is designed to answer is whether it can be the foundation for a $10B+ infrastructure company, or whether it gets commoditised as the hyperscalers build equivalent abstractions into their existing platforms. The bet is that developer loyalty, performance compounding, and the head start in serverless GPU architecture create a moat that is harder to replicate than it looks from the outside.
Open roles at Modal Labs
6 positions we're tracking. Roles are re-checked daily and removed when filled.
Infrastructure / Systems Engineers
First seen 4 months ago
GPU / CUDA Optimization Engineers
First seen 4 months ago
Developer Advocates
First seen 4 months ago
Enterprise Sales
First seen 4 months ago
Developer focused PM
First seen 4 months ago
Site Reliability Engineers
First seen 4 months ago
Hiring outlook
Valuation 2.3x in 5 months, $50M ARR
Working at Modal Labs
Founded in , Modal Labs is AI Inference Infrastructure. They're now people. For a AI company this size, the reality is opportunity to shape your role based on the company stage.
The majority of roles are in San Francisco.
Frequently asked questions
How many jobs does Modal Labs have open?
As of February 2025, Modal Labs has 6 open positions.
Does Modal Labs hire remotely?
Modal Labs doesn't have remote openings at the moment. All roles are in San Francisco.
What roles is Modal Labs hiring for?
Modal Labs is hiring across Engineering, Sales. The most recent opening is Infrastructure / Systems Engineers.
How do I apply for a job at Modal Labs?
Use the apply links above, or check our guide to getting hired at Modal Labs.
Where is Modal Labs based?
Modal Labs is headquartered in San Francisco.
Get Modal Labs roles before they're posted
We find AI roles at companies like Modal Labs before the job boards do. Fewer applicants, better odds.
Get early access →