Sovereign AI is only as strong as the data that powers it

By Varun Babbar, VP & India MD, Qlik 

India is racing to build a sovereign artificial intelligence stack, and the momentum is real. Through the IndiaAI Mission, the government is stitching together compute capacity, indigenous foundation models, talent pipelines and a governance framework meant to give the country control over its own AI future. AI Kosh, the national data and compute repository, has already brought together more than 7,500 datasets and 273 AI models across 20 sectors. That is a genuine achievement, and it deserves recognition.

But there is a quieter problem sitting underneath all this ambition: making data available is not the same as making it usable. Data that powers AI systems in India, as in most large economies, remains scattered across departments, companies and legacy systems, with gaps in quality, accessibility and oversight. If data readiness does not keep pace with investment in chips and models, India risks ending up with powerful AI infrastructure that has no trustworthy foundation underneath it: capable of impressive demonstrations, but not dependable, explainable outcomes.

This is the part of the sovereignty debate that gets the least attention. Sovereign AI is often discussed as a matter of geography: where servers sit, where models are trained, who owns the compute. Those questions matter. But sovereignty is hollow if it is only about location. It also has to mean the ability to understand, govern, connect and use data effectively, so that control over information actually converts into value from AI. A country can hold all its data within its own borders and still not know where that data lives, who owns it or whether it can be trusted.

Consider what the IndiaAI Mission is already building. It is not narrowly focused on model development. It spans compute, datasets, innovation, talent and responsible AI as a connected ecosystem. The government’s National Data Governance Framework, approved this year, treats data as a strategic factor of production, on par with capital or labour, and pushes for common standards, shared metadata and APIs that make secure data sharing possible. Platforms such as AI Kosh and API Setu are the building blocks meant to make Indian data more discoverable across applications. All of this is necessary. None of it is sufficient.

That fragmentation becomes an AI problem very quickly. AI systems are only as good as the data that trains, grounds and operates them. Inconsistent definitions, incomplete records and disconnected systems can produce unreliable outputs even when the underlying compute is world class and the model is state of the art. India’s scale makes this harder, not easier. A country with hundreds of languages, wide variation in economic conditions and a patchwork of regulatory regimes needs data that reflects those conditions if its AI systems are going to be useful rather than merely impressive in a demo.

This is where governance, lineage and integration stop being back-office concerns and start functioning as core AI infrastructure, on the same tier as GPUs and model weights. Governance has to move past compliance checklists and establish real ownership, quality standards and access controls. Lineage matters because it shows where data originates, how it changes as it moves through systems, and which AI outputs depend on it, which is the only way to build auditability into a system that regulators and citizens are right to be sceptical of. Integration matters because it connects the departments and legacy systems that otherwise keep useful information locked away from each other. 

In regulated sectors such as banking, financial services, insurance and health care, the ability to trace an AI decision back to its governed source data will decide whether these systems are trusted with real responsibility or kept at arm’s length.

None of this argues against India’s investment in compute and models. It argues for treating data readiness as an equal priority: identifying which datasets actually matter, assigning clear ownership, eliminating duplication, connecting silos and putting real controls around sensitive information. For enterprises, the priority should similarly shift from accumulating more data to making the data they already hold discoverable, trusted and accessible to the systems authorised to use it.

Done well, this does not require India, or any Indian organisation, to wall itself off. Sovereignty and interoperability are not opposites. Organisations can retain control over sensitive data while still connecting it, securely, across approved systems, which is exactly the kind of foundation that lets AI pilots graduate into deployments with real accountability behind them.

India’s sovereign AI question is no longer whether the country can build the infrastructure. It clearly can. The real question is whether that infrastructure will produce intelligence that is reliable, locally relevant and trusted. Compute supplies the capacity. Models supply the intelligence. But it is trusted, governed data that supplies the context, and without it, sovereignty is just a flag planted on top of someone else’s foundation.

AIQlikSovereign AI
Comments (0)
Add Comment