'Public Does Not Mean Permissionless': The DPDP Act Loophole That Just Became India's Biggest AI Fight
Nasscom reopened the DPDP Act's Section 3(c)(ii) 'publicly available data' exemption on July 28. Here's why AI firms can't rely on it to scrape.
#'Public Does Not Mean Permissionless': The DPDP Act Loophole That Just Became India's Biggest AI Fight
If your company trains AI models on data scraped from the open web, the single most important sentence in Indian privacy law is nine words long — and nobody agrees on what it means. On 28 July 2026, that disagreement moved back to centre stage: an industry discussion surfaced through Nasscom, the body that speaks for India's IT and AI sector, arguing that the DPDP Act's "publicly available data" exemption is now the sharpest unresolved conflict between personal-data rights and AI development in the country. It landed the same day compliance analysts flagged that India's data-privacy rules are entering their enforcement phases, raising the stakes for every data-heavy business — foreign financial firms included. The short version for Indian businesses: the exemption you think lets you scrape public data may not, and betting your model pipeline on it is a bet the government has already signalled it will contest.
This is not an academic quarrel. It decides whether an Indian AI startup can legally build a model on Indian-language social media posts, whether a bank can enrich customer profiles with scraped public data, and whether a Significant Data Fiduciary's training corpus is a compliance asset or a ₹250 crore liability. The Digital Personal Data Protection Act, 2023 is now in force, the DPDP Rules 2025 are notified, and the runway to full enforcement is short. The one question the law left dangerously vague is the one AI teams most need answered.
#The nine words at the centre of the fight
Section 3(c)(ii) of the DPDP Act says the law does not apply to personal data "that is made or caused to be made publicly available by the Data Principal to whom such personal data relates," or by any other person under a legal obligation to publish it. On a plain reading, that is sweeping. A LinkedIn profile you set to public, a tweet, a self-published blog, parliamentary transcripts, court records, and a company's statutory filings all appear to fall entirely outside the Act. No consent, no purpose limitation, no notice — the protections simply switch off.
For AI developers, that reading is a lifeline. Modern models are trained on enormous corpora scraped from public webpages, social platforms, reviews and forums, and a huge share of that text contains personal data — names, opinions, locations, relationships. If Section 3(c)(ii) means what it appears to say, then a dataset assembled from public sources sits beyond the DPDP Act's reach, and Indian firms can compete with global labs without licensing every byte.
The problem is that the government does not read it that way.
#The government's counter-reading: "public does not mean permissionless"
Back in August 2024, the Minister of State for Electronics and IT told the Rajya Sabha that scraping publicly available user data remains subject to the Information Technology Act, the IT Rules, and the DPDP Act — including its consent and transparency obligations. In other words, the ministry's position is that making data public does not strip an individual of their data-protection rights, and a scraper still needs a lawful footing. As one widely-cited IAPP analysis put it bluntly: "public does not mean permissionless."
That leaves companies staring at a genuine contradiction. The statute contains a categorical carve-out; the executive insists the carve-out is narrow and that consent obligations survive. Legal commentators have called this exactly what it is — a fundamental contradiction between the text of Section 3(c)(ii), the government's stated view, and the overlay of the IT Act. When the Data Protection Board becomes fully operational and a complaint about scraped data crosses its desk, which reading wins is anyone's guess. That uncertainty is itself the compliance risk.
#Why the exemption is narrower than it looks
Even taken at face value, Section 3(c)(ii) is riddled with limits that AI teams routinely overlook.
"By the Data Principal" is doing heavy lifting. The exemption applies where the individual themselves made the data public — or someone under a legal duty to publish it did. Data that a third party reposted, leaked, or re-shared without the original person's agency arguably does not qualify. Consider the mid-July 2026 revelations that over 1,000 databases of Indian students' exam records were being sold online, and that the government's own UMANG platform stored Aadhaar numbers in plaintext. That data is "publicly available" in the crudest sense — anyone can buy or grab it — but none of it was made public by the data principal. Scraping it does not fall inside Section 3(c)(ii); it falls inside the definition of a breach.
Platform defaults are not "meaningful choice." The statute never defines what "made publicly available" means. Does a social account that is public by default — because the user never changed a setting they may not have understood — count as the individual choosing to publish? Legal scholars argue the exemption should turn on meaningful agency, not on a toggle the platform pre-set. Until the Board clarifies this, treating every public-by-default profile as fair game is an aggressive interpretation, not a safe one.
Hybrid pipelines trigger "re-entry." Most real AI workflows do not process pristine public data in isolation. They blend it with user records, purchased datasets, feedback logs and enrichment sources. The moment exempt public data is combined with in-scope personal data, the exemption's factual conditions stop being satisfied for the combined set, and the Act re-applies. A "public data" defence collapses the instant your pipeline mixes sources — which nearly all of them do.
The research exemption doesn't rescue you either. Section 17(2)(b) offers a conditional carve-out for processing for research, archiving or statistical purposes, provided the data is not used to make decisions "specific to a Data Principal." Large-scale model training strains that limit: a trained model produces persistent, cross-context outputs that can affect identifiable individuals, even if no single training step is framed as a decision about them. Leaning on the research exemption for commercial model-building is a stretch most counsel will not sign off on.
#Industry's ask: exempt AI training outright
This is why India's tech lobby has been pushing hard. The Internet and Mobile Association of India (IAMAI) has written to MeitY seeking a clear exemption for companies training AI models, arguing that the ambiguity in Section 3(c)(ii) will raise entry barriers, inflate development costs, and hand an advantage to the few firms rich enough to license proprietary datasets. IAMAI wants the government either to amend the Act or to use its power under Section 17(5) — which lets the Centre exempt classes of data fiduciaries by notification — to take AI training on publicly available data out of scope.
Nasscom's 28 July framing widens the lens beyond privacy, tying it to copyright and trade-secret exposure: the same scraped corpus that raises DPDP questions can simultaneously infringe copyright and absorb confidential business information. The industry's message to MeitY is that AI developers face a three-front legal conflict on a single dataset, and the government's silence on the publicly-available-data question is the most urgent gap to close before enforcement bites.
The counter-argument is equally serious. Handing AI developers a blanket pass would gut the Act's protections precisely where modern harm concentrates — profiling, inference and re-identification built from "public" scraps. The Supreme Court's foundational Puttaswamy reasoning cuts against a permissive reading: privacy interests do not evaporate merely because information is accessible in public. A middle path — time-bound relief for startups, mandatory disclosure of training-data categories, and clear rules on what "public" actually means — is what several analysts now advocate.
#GDPR shows how different India's choice is
It helps to see how unusual India's design is. Under the EU's GDPR, there is no blanket exclusion for publicly available personal data. A controller scraping public profiles still needs a lawful basis, and must honour purpose limitation, data minimisation and transparency. Public accessibility is a factor in the risk and fairness assessment — it never removes the data from the regulation. That is why European regulators have been able to act against indiscriminate scrapers even when the source data was technically public.
India took the opposite structural route. Section 3(c)(ii) is an on/off threshold: qualify, and the entire Act drops away with no residual obligations. That makes India's text simultaneously more permissive than the GDPR (a true exemption exists) and more precarious (its boundaries are undefined, and the executive is trying to narrow it through interpretation rather than drafting). For multinationals running a single global data pipeline, this GDPR vs DPDP Act divergence is not a footnote — a scraping practice engineered to satisfy Europe's lawful-basis logic may not map cleanly onto India's binary exemption, and vice versa.
#What Indian businesses should do before the meter switches on
The DPDP Rules were notified on 13 November 2025; the consent-manager framework is expected to go live around 13 November 2026; and the substantive obligations, backed by penalties up to ₹250 crore, become fully enforceable by 13 May 2027. The Data Protection Board has been constituted and its members appointed. The window to fix data-sourcing practices while mistakes are still cheap is closing. Concretely:
- Map the provenance of every training dataset. For each source, record who made the data public, how (deliberate publication vs platform default vs leak), and when. If you cannot show the data principal themselves made it public, do not assume Section 3(c)(ii) covers you.
- Segregate exempt and non-exempt data. The "re-entry" problem means blended pipelines forfeit the exemption. Keep genuinely-public-by-choice data on a separate, documented track from user, purchased or enriched data, and apply full DPDP controls to the latter.
- Stop treating leaked or breached data as "public." Buying a scraped exam-records database or ingesting a plaintext Aadhaar dump is not exemption territory — it is participation in a breach. Build source-hygiene checks that reject data of unlawful origin.
- Don't bank on a Section 17(5) rescue. IAMAI's requested AI exemption may never come, or may arrive with conditions. Plan for the world where consent and transparency obligations do apply to scraped personal data, and treat any future carve-out as upside, not baseline.
- Prepare SDF-grade documentation now. If your data volumes point toward Significant Data Fiduciary designation, your DPIAs and audit trail must already answer: what public data do we train on, on what legal footing, and how do we honour erasure and correction requests for individuals in that corpus?
Our /resources library includes a data-provenance and lawful-basis template built for exactly this Section 3(c)(ii) grey zone, and our consent-manager guide covers how to capture verifiable consent where the exemption does not reach — the safer footing for any dataset you cannot prove was made public by the individual.
#The bottom line
The DPDP Act handed AI developers what looks like a gift — a categorical exemption for publicly available personal data — and then left it undefined enough to be a trap. The government reads it narrowly, the industry reads it broadly, and the text supports an argument either way. Nasscom putting this conflict back on the table on 28 July is a signal that the sector wants certainty before the Board starts enforcing, not after.
For Indian businesses the safe posture is the conservative one: public does not mean permissionless. Treat scraped personal data as in-scope until you can prove otherwise, document provenance obsessively, and never let a breach masquerade as a "public source." The companies that build that discipline into their data pipelines now — rather than waiting for MeitY to settle the argument — are the ones that will still be training models in India when the enforcement meter finally switches on in 2027.
Sources: Nasscom — Privacy, Copyright, and Trade Secret: the conflict between personal data rights and AI training; DPDP News tracker; IAPP — Scraping public data in India; Law School Policy Review — Publicly Available Data under the DPDP Act; MediaNama — IAMAI seeks AI training exemption; MediaNama — student exam data sold online; MediaNama — UMANG plaintext Aadhaar exposure.