Structured vs Unstructured Data: Why This Matters More Than Ever, and What to Do About It
(Updated from an article published Jan. 1, 2026 by ALM)
Recent reporting keeps confirming the same pattern: large-scale supply chain breaches involving core SaaS platforms are no longer hypothetical, they’re routine.
In June 2026, a breach at Klue, a B2B sales intelligence platform, gave attackers OAuth access into the Salesforce environments of roughly 200 downstream customers, exposing CRM records, business contacts, and sales data across the chain. Around the same time, a separate wave of Oracle breaches hit both Oracle Cloud Infrastructure and Oracle E-Business Suite, exposing everything from patient records to the financial ledgers, vendor contracts, and payroll data that large enterprises run their operations on, with confirmed victims spanning multiple industries.
According to IBM’s 2026 Cost of a Data Breach Report, the global average cost of a breach climbed to $4.99 million this year, a 12 percent jump and a new record, driven largely by higher detection, escalation, and lost-business costs. Breaches tied to AI-enabled attacks ran even higher, averaging $6 million each.
These incidents are not about stolen documents, but rather exhaustive database rows. A PDF letter might expose one person at a time, but a CRM table can expose hundreds of thousands of people in one row set. That is a different problem entirely.
This shift is important because most breach response data mining strategies have been built on the world of unstructured documents. The majority of the world’s personal data is now sitting, however, inside relational architecture.
Two universes of data
Unstructured data:
- Narrative files
- Email bodies, attachments, PDFs, call notes, scanned images
- Interpretation is semantic
- The reviewer reads context and identifies PII inside sentences and paragraphs
Structured data:
- Relational systems
- Rows, columns, key value relationships, schema logic, parent child table mapping
- Interpretation is relational
- The reviewer identifies which fields represent identity attributes and what specific records for real individuals were exposed
These are not the same skill set. An unstructured review mindset requires reading. A structured review mindset demands data model reasoning.
Structured data requires a different workflow
Based on experience with recent Salesforce and Oracle structured data breaches, I would suggest a six-step workflow tailored to the precise nature and corresponding problems with structured data review. This workflow is not your typical data mining standard of loading data into an eDiscovery tool, running a combination of search terms, machine learning, and/or AI to identify potentially sensitive documents. Instead, a structured data review is surgical, allows the data to remain in its original format, and requires a different type of reviewer with a technical, data model reasoning workflow.
Here are the six basic steps for the structured data workflow:
This process creates a defensible chain of logic and avoids the most common failure mode in database incidents: assuming column names equal plain language meaning.
What recent incidents teach us about response
The Salesloft/Drift and Gainsight campaigns are, in effect, a live case study in how a structured data incident actually unfolds, and what separates a contained breach from a spreading one. A few lessons stand out:
Revoke fast, revoke broadly.
Salesloft and Salesforce detected the Drift compromise on August 19 and had revoked every active Drift token and pulled the integration from the AppExchange by the next day. That speed is the standard in which to hold vendors and internal teams to, not the exception.
Assume the exposure doesn't end with the first vendor
The Gainsight incident followed roughly three months after Salesloft/Drift, when attackers used credentials harvested from the earlier breach to compromise a second, unrelated vendor and pivot into roughly 285 more Salesforce instances. A token-theft event should trigger a review of everything those tokens, and anything derived from them, could have touched, not just the system where the compromise was first detected.
Hunt for secondary exposure inside the data itself
Support cases and CRM records frequently contain embedded credentials, such as API keys, passwords, or cloud tokens. A structured data review needs to search the exposed records for these secrets specifically and rotate them, not just reset the accounts that were directly compromised.
Build a third-party incident response plan, not only an internal one
In both campaigns, the initial compromise happened at a vendor, which means an organization’s response time was gated by that vendor’s detection and disclosure speed. Rapid-notification SLAs with SaaS and integration vendors should be a contractual requirement, not an assumption.
Plan for extortion, not just data theft
The threat actors behind these campaigns have leaned increasingly on extortion rather than simple data theft. An incident response plan needs a legal and negotiation track running alongside the technical containment work, not bolted on afterward.
Building the capability before you need it
Relational data literacy cannot be improvised in the middle of a breach. It has to be built into an organization’s incident response capability well before an incident occurs:
- Maintain living schema documentation for every system that holds regulated data, treated as an incident response prerequisite rather than an engineering artifact.
- Enable and retain detailed query and export logging on core relational systems, so that a breach can be reconstructed at the query level, not just the access level.
- Classify data at the column level in advance, so that scope statements after a breach can be produced quickly and defensibly.
- Ensure incident response teams include SQL-literate analysts, not only reviewers trained in document-style forensic review.
- Inventory third-party OAuth and API access with the same rigor as privileged internal accounts, including exactly what tables and joins each integration’s token can reach.
The common thread across the last year’s largest breaches was never a flaw in the core platform. It was a trusted integration that was never scoped tightly enough to contain the damage once it was, inevitably, compromised.
Why this is a critical pivot moment
Supply chain breaches are no longer just hitting endpoints. Threat actors are targeting the systems at the heart of running a business. When the data compromised is not an email archive but the CRM or ERP system that contains every customer record and every linked attribute across dozens of dependent tables, the breach response lens must shift. Structured data review is database forensics that is now a critical part of incident response.
Given how the threat landscape is evolving, structured data expertise is no longer a niche specialty. It should now be a core capability for modern incident response. The pattern across Salesloft, Gainsight, and the campaigns that followed them is not that attackers found a new vulnerability class. It’s that they found an interpretive gap, between how relational systems actually work and how breach response teams have traditionally been trained to think about exposure.
The next wave of breach response will not be won by who can read documents faster (AI will probably handle that soon enough anyway). It will be won by who can close that interpretive gap: interpreting relational data structures accurately, in combination with the data breach response expertise built in the unstructured data world. Closing that gap is now a baseline requirement, not a specialization.
CIPP/US – Vice President, Cyber Solutions
lntegreon
About the author
Megan Silverman is Vice President, Cyber Solutions at Integreon and works with clients to deliver customized approaches to handle data mining and notification list development for large-scale complex breaches. Megan has deep expertise in litigation, privacy, and cyber incident response. She is a certified privacy professional, earned her JD from the University of Chicago, and her Environmental Law LLM at Lewis & Clark Law School. Megan is a member of the New York bar.