Social Media Scraping Under DPDP: What Indian Marketers and Sales Teams Must Know
Scraping profiles, contact lists and post data from social platforms is standard practice in marketing and sales. The DPDP Act's publicly available data exclusion is narrower than most teams assume. Here is where the line actually sits.
On This Page
The Practice Everyone Uses and Nobody Documents
Social media scraping is a default behaviour in Indian marketing and sales. Teams pull profile lists from LinkedIn, contact details from directories, and post data from X, Instagram and Facebook. They build lead lists, enrichment databases and lookalike audiences from what they find. The tools that do this are widely advertised, cheap, and increasingly automated.
Under the DPDP Act, that activity has a compliance cost that most teams have not priced in. The data being scraped is personal data. The Act does not distinguish between data collected through a form and data collected through a crawler. The obligations attach to the processing, not to the method of collection.
The scraping industry’s standard defence is a single phrase: the data is publicly available. The Act does contain an exclusion for publicly available data. It is narrower than the industry assumes, and the difference is where the entire compliance question sits.
What Section 3(c)(ii) Actually Excludes
Section 3(c)(ii) states that the Act does not apply to personal data made or caused to be made publicly available by the Data Principal herself, or by another person under a legal obligation to publish it. Data that qualifies is outside the Act. No consent requirement, no notice, no erasure obligation.
Read the condition again. The exclusion turns on who made the data available, not on whether the data can be found in public. Two consequences follow:
- What the principal published is excluded. A founder who writes company updates on a public LinkedIn profile has made that data publicly available within the meaning of Section 3(c)(ii). The Act does not apply to it. Statutory disclosures sit in the same category: a director’s name on the MCA register is published under legal obligation.
- Almost everything else in a scraped database is not excluded. An email address inferred from a name-and-domain pattern was never made available by the principal. A phone number assembled from fragments across sources was not. Enrichment data purchased from a vendor, inferences drawn from post history, and profiles compiled from what colleagues or platforms disclosed all fail the condition. The Act applies to that data in full.
A scraped lead database is almost always a mixture of both categories. The fiduciary carries the burden of knowing which records are which. No court or Board order has yet drawn the boundary of Section 3(c)(ii) for scraped social data, so the defensible position is the documented one: record, per source, who made the data available and for what.
Where In-Scope Scraped Data Creates Obligations
Lawful basis. Section 4(1) permits processing on exactly two grounds: consent under Section 6, or the legitimate uses listed in Section 7. Section 7 is a closed list. Voluntary provision by the principal, state functions, medical emergencies, employment purposes and similar grounds. Direct marketing is not among them, and the Act contains no legitimate-interest ground comparable to the GDPR. Marketing outreach to in-scope scraped contacts therefore requires consent, and a list built without it is the single most common exposure in this area.
Security safeguards. A scraped database is personal data in your possession. Section 8(5) requires reasonable security safeguards to prevent a breach. Scraped databases concentrate contact data, accumulate stale and duplicated records, and are prime targets for theft. The ₹250 crore cap in the Schedule attaches to exactly this failure.
Breach notification. If a scraped database is compromised, Section 8(6) requires intimation to the Data Protection Board and each affected Data Principal. Rule 7 of the DPDP Rules 2025 sets the operational detail: affected principals are informed without delay, and the Board receives a detailed report within 72 hours of awareness. Rule 7 becomes enforceable on 13 May 2027. Teams that treat scraped lists as operational assets rather than regulated data rarely have a breach plan that covers them.
Erasure and correction. Data principals hold correction and erasure rights under Chapter III of the Act. A principal who appears in your scraped lead database has the same erasure rights as one who submitted a form on your website. Your response mechanism must reach both.
Vendor accountability. Most scraping is outsourced. A vendor processing personal data on your behalf is a processor under the Act, and you remain the Data Fiduciary. The contract must cover purpose, safeguards and breach notification support. Vendors that train models on scraped data, or that hold it outside India, add purpose and cross-border exposures that belong in your transfer documentation.
What Teams Should Do Now
- Inventory the scraping. Document every tool, script and vendor that collects personal data from social platforms. Name the data fields, the sources, and the purpose.
- Classify against Section 3(c)(ii). For each source, record who made the data available. Principal-published profile data sits outside the Act. Inferred, assembled and enriched data sits inside it. The classification is the compliance position.
- Map the lawful basis for in-scope data. For each use, record whether a consent record exists or a Section 7 ground genuinely applies. Where neither does, stop the activity or build a consent mechanism.
- Review marketing lists. Lists built from inferred or enriched contact data without consent are an exposure. Replace them with opt-in acquisition, or obtain a legal review of the basis you believe applies.
- Add scraped databases to your security and breach plan. They are personal data stores. They belong in your Section 8(5) safeguards, your breach response plan and your erasure workflow.
- Audit your vendors. Confirm which scraping vendors are processors, check the contract covers the Act’s requirements, and document cross-border flows for vendors outside India.
The Positioning Trap
There is a temptation to read this article as a reason to avoid scraping entirely. That is not the point. Scraping principal-published data for research, competitor monitoring and market intelligence sits outside the Act by operation of Section 3(c)(ii). The point is classification and record. The same activity is either a defensible process or a penalty exposure depending entirely on whether the analysis is documented.
That is the DPDP pattern across every area of compliance: the underlying activity is rarely prohibited, but the undocumented version is the expensive one. Source classification, consent records and erasure workflows turn the same behaviour into evidence.
Run the free DPDP gap assessment to see how your data inventory, consent records and breach readiness hold up.
Sources
- Digital Personal Data Protection Act, 2023: Section 3(c)(ii) (publicly available data exclusion), Section 4(1) (lawful grounds), Section 6 (consent), Section 7 (certain legitimate uses), Sections 8(5) and 8(6) (security safeguards and breach intimation), Chapter III (data principal rights), the Schedule (penalties)
- Digital Personal Data Protection Rules, 2025 (G.S.R. 846(E), notified 13 November 2025): Rule 7 (breach intimation), commencement schedule under Rule 1
- ConsentOS Learn Hub: What Is the DPDP Act 2023?
Frequently asked questions
Is scraping publicly available social media data allowed under the DPDP Act?
Sometimes, and the boundary is precise. Section 3(c)(ii) excludes from the Act personal data that the Data Principal herself made publicly available, or that another person was legally obliged to publish. Data inside that exclusion is outside the Act. The problem is that most scraped datasets do not qualify: inferred email addresses, phone numbers assembled from fragments, enrichment data, and profiles compiled from what other people posted were not made available by the principal, so the Act applies to them in full. A scraped lead database is almost always a mixture, and the burden of separating excluded data from regulated data sits with the fiduciary.
Does the DPDP Act require consent to send marketing emails to scraped contacts?
For data within the Act's scope, yes. Section 4(1) permits processing only on consent under Section 6 or for the legitimate uses listed in Section 7. The Section 7 list is closed: voluntary provision, state functions, medical emergencies, employment purposes and similar grounds. Direct marketing is not on it, and the Act has no legitimate-interest ground comparable to the GDPR. A contact whose email address was inferred or assembled by a scraping tool never made that data publicly available, so the Section 3(c)(ii) exclusion does not rescue the list. Without a consent record, the list is a compliance exposure, not a data asset.
What penalties apply to non-compliant scraping under DPDP?
The Schedule to the Act sets the caps: up to ₹250 crore for failure to take reasonable security safeguards to prevent a personal data breach, and up to ₹50 crore under the residual tier for other contraventions, including processing without a lawful basis. The Data Protection Board of India adjudicates penalties. The substantive Rules obligations become enforceable on 13 May 2027, which makes the current period a preparation window, not a grace period.
Are scraping tools data processors under DPDP?
If a scraping tool or vendor processes personal data on your behalf, it is a processor under the Act and you remain the Data Fiduciary. Accountability stays with you: the contract must cover purpose, security safeguards and breach notification support. Vendors that train models on scraped data, or that store data outside India, add a purpose and cross-border exposure you must document.
Know where you stand on DPDP compliance
Run the free DPDP Gap Assessment for a gap report scored against your DPDP Act 2023 obligations, work through the 26-point compliance checklist, or model your penalty exposure.
Enforcement milestones, rule notifications, and deadline analysis.
One email when it matters, no more.
Resources
Continue Reading
Related DPDP Act 2023 guidance from the ConsentOS knowledge base.
DPDP Compliance Checklist: 43 Controls for Indian Businesses (2026)
Audit your DPDP Act 2023 posture against 43 controls, then sequence remediation across five months to the November 2026 Consent Manager deadline.
10 min read
Regulatory UpdatesDPDP Penalties: ₹250 Crore Risk and Enforcement Tiers in India
A breakdown of every penalty provision in the DPDP Act 2023. Understand the financial exposure, the enforcement mechanism, and what triggers each penalty tier.
7 min read
Data Principal RightsDPDP Act 2023: All 8 Data Principal Rights with Templates (India)
Access, correction, erasure, grievance, and nominee rights under the DPDP Act 2023: the response deadlines a Data Fiduciary must meet, with ready-to-use request templates.
7 min read