CoRover.ai, the company behind the BharatGPT family of sovereign foundational AI models, has released a new set of benchmark results that highlights the performance and efficiency of smaller AI models built for India. The latest evaluation shows that BharatGPT-Instruct, a model with approximately 2 billion active parameters, outperformed Sarvam 30B, a model roughly 15 times its size, across key reasoning and fairness benchmarks. BharatGPT-Instruct also led Qwen 1.7B across all ten regional languages tested on Multilingual MMLU.
The results reinforce CoRover’s positioning that the future of sovereign AI should not be measured simply by model size or parameter count. Across ARC-Easy, ARC-Challenge, WinoGrande, PIQA and CrowS-Pairs, BharatGPT-Instruct recorded scores of 83.21%, 53.84%, 68.67%, 78.51% and 68.28%, respectively. On each of these benchmarks, the model outperformed both Qwen 1.7B and Sarvam 30B, demonstrating that a significantly smaller model can deliver competitive, and in several areas stronger, reasoning and fairness performance.
Ankush Sabharwal, Founder and CEO, CoRover.ai, said, “Scale is not the same as intelligence. These results demonstrate that when AI is designed around the problems, languages and requirements of a country, a smaller model can deliver remarkably strong performance. For us, sovereign AI is not about building the biggest model; it is about building models that are efficient, relevant, trustworthy and capable of solving real-world problems at scale. The next phase of AI will be defined by how intelligently we build, not simply by how large we build.”
A key highlight of the evaluation is BharatGPT-Instruct’s multilingual performance. Tested across Hindi, Bengali, Marathi, Telugu, Gujarati, Malayalam, Punjabi, Tamil, Odia and Kannada, the model led both Qwen 1.7B and Sarvam 30B across all ten languages on Multilingual MMLU. The results underline CoRover’s focus on developing AI that can understand and reason across India’s linguistic diversity, rather than relying primarily on English-centric capabilities.
For enterprise deployments, CoRover has placed particular emphasis on retrieval-augmented generation, or RAG, where the quality of an AI system depends on how faithfully it responds using an organisation’s own knowledge base. In the company’s evaluation, BharatGPT-Instruct achieved 100% on Faithfulness, ahead of GPT-4o-mini at 95% and Qwen 3 1.7B at 95.51%. It also achieved 100% Top-K Accuracy, matching the other models evaluated on that metric.
Ankush Sabharwal added, “Enterprise AI needs to go beyond generating fluent responses. Organisations need systems that are grounded in their own information, operate within clearly defined governance frameworks and can handle sensitive data responsibly. Our RAG results are particularly significant because they reflect the kind of performance that matters in real enterprise environments. We believe sovereign AI must bring together intelligence, multilingual capability, security, data control and accountability.”
Alongside model performance, CoRover has highlighted data governance and security as core components of its sovereign AI approach. The company said client data is processed but not owned by CoRover, while datasets entering BharatGPT training undergo source and licence verification, contractual and right-to-train checks, PII filtering and anonymisation, as well as quality and deduplication checks. For sensitive institutional and regulated partners, the company follows a no-training, deployment-only policy.
CoRover is also developing domain-specific versions of BharatGPT designed around the workflows and language requirements of individual sectors. PortGPT, built for the shipping ports domain, is already live, while additional domain-specific models are in development. The company said its broader platform has reached more than 1.8 billion lives, 65 million-plus monthly active users and 95,000-plus enterprises, developers and researchers across more than 100 languages and 20-plus channels.
Importantly, CoRover has also disclosed where BharatGPT currently has room for improvement. On the GSM8K mathematics reasoning benchmark, BharatGPT-Instruct scored 23.50%, compared with 68% for Qwen 1.7B and 70.60% for Sarvam 30B. The company said publishing these gaps is deliberate, allowing future development to focus on areas where the model needs to improve.

