tech · 2026-07-31
Sarvam AI Bets Big on Trillion-Parameter Model

Photo: Solutionservice21 / Wikimedia (CC BY-SA 4.0)
Sarvam AI plans a trillion-plus parameter foundation model, ready in ~6 months, plus new coding and voice AI tools.Its 105Bn-param model costs $0.80/million tokens, over 5x cheaper than GPT-5.4 Mini and 11x cheaper than Gemini 3.5 Flash.Comes after Sarvam's $234Mn first close of a $300Mn Series B round, valuing it at $1.5Bn.
Does building models from scratch make sense?
Building from scratch lets Sarvam optimize for Indian languages and agentic tasks, unlike using someone else's foundation model, said co-founder Pratyush Kumar, who wants India to compete in coding, cybersecurity and science.
What does 'agentic' benchmark performance test?
Agentic benchmarks test whether a model can plan multi-step tasks and use tools, not just answer questions, which matters for coding and cybersecurity applications Sarvam is targeting with its trillion-plus parameter model launching in six months.
Why open a San Francisco office to hire?
Kumar said the goal is a 'strong conduit of talent back to India,' tapping the world's largest AI talent pool through a San Francisco office while keeping development focused on Indian context.
What does the $234Mn Series B money actually fund?
The Series B funds research on Sarvam's next frontier model for agentic, coding and cybersecurity uses, plus compute access to expand deployment across key verticals, with a $234 million first close valuing the company at $1.5 billion post-money.
Can Sarvam's pricing edge survive at scale?
Sarvam's 105Bn-param model already prices at $0.80/million tokens versus $4.50 for GPT-5.4 Mini, a 5x+ gap that widens further against Gemini 3.5 Flash's $9 rate. Whether that gap holds as models scale to a trillion parameters is untested.
How does token pricing work for AI models?
Token pricing charges per million words/characters processed, blending input and output costs. Sarvam's 105-billion-parameter model at $0.80 rate versus GPT-5.4 Mini's $4.50 shows how domestic training and inference can undercut global providers per unit.
What happens if compute costs rise at scale?
Larger models need more chips and power, so Sarvam's plan to expand compute infrastructure is central to whether it can keep prices low even as parameter count multiplies toward a trillion, particularly for coding, cybersecurity, and simulation tasks.
Why does data residency matter for Indian cos?
Hosting models in India lets enterprises and government agencies keep sensitive data within domestic borders, addressing compliance concerns that using foreign-hosted models like GPT or Gemini would raise. This supports Sarvam's goal of producing the largest share of AI tokens consumed in India domestically.
Who beyond Indian developers gains from this push?
Enterprises, developers and govt bodies in India get access to frontier AI hosted domestically, per Sarvam, with data staying in-country rather than routed through foreign clouds.
Who decides which Indian languages get priority?
Sarvam said its 105-billion-parameter model outperforms peers specifically on Indian-language benchmarks, suggesting language coverage is chosen based on where it can beat global competitors, not just broad market size.
How do voice tools like Bulbul V4 fit the strategy
Bulbul V4, a multilingual text-to-speech model, and Saaras V4, a speech recognition model, extend Sarvam's stack beyond text into voice, targeting India's large non-English-speaking user base. The company also launched Sarvam Code, an AI coding assistant.
What does 'largest share of AI tokens' mean?
Kumar said Sarvam wants to "produce the largest share of AI tokens consumed in India," meaning most AI queries and outputs generated in India would run on Sarvam's infrastructure rather than foreign providers, with models hosted domestically and priced at $0.80 per million blended tokens.
Source: thehindubusinessline.com