Impact, Compliance, and Countermeasures in Relation to Data Breaches in Publicly Traded U.S. Companies (Future Internet, 2024)
Rodrigues, Serrano, Vergara, Albuquerque and Amvame-Nze, Future Internet 16(6):201, 5 June 2024 (MDPI, CC-BY). An empirical pass over 506 breaches at 274 NYSE/NASDAQ-listed companies, taken from Rosati and Lynn’s filtered slice of the Privacy Rights Clearinghouse repository and augmented with company data. Ingested via the research pass against the spoke’s compliance edge.
Why it was hunted
The edge asks whether compliance correlates with not getting breached. This is the closest thing to a dataset the search produced. It does not answer that question, and the reason is worth holding.
What the data says
- 506 breaches, 274 companies, 2005 to March 2015. PRC’s records come from state Attorneys General and HHS, so the set is reported breaches, not all breaches — the denominator problem is inherited.
- Roughly 1.074 × 10⁹ records in total, about three times the 2013 US population.
- Two breach types are half of everything: PORT (lost or stolen portable devices) at 139 events and HACK (malicious outsiders) at 118, together 50.79%. PORT dominates the early years — 43 events in 2006, close to four times the next type.
- Human-factor breaches decline over the window (the authors group INSD, PHYS, PORT, STAT and DISC), agreeing with Hammouchi et al. on direction.
- Finance is the most-breached sector, and four of the ten most-breached companies are financial. This contradicts Hammouchi et al., who worked the same PRC source over 2005–2019 and found healthcare and the business/technology grouping on top. The disagreement is in the paper and is not resolved by it.
- 185 of 506 breaches were preceded by a company announcement, most often an earnings release.
The compliance half, and why the edge stays open
Section 5.1 maps each breach to the standard that would govern it — SOX, HIPAA, GLBA, PCI-DSS — and counts breaches per law. That is coverage of a regulation, not compliance with it: no company in the dataset carries an audit status, an attestation, or any variable saying it met the standard. Nothing here can compare a compliant estate with a non-compliant one.
The timing makes it worse in a way the authors state plainly. The dataset runs 2005–2015; the first US state data protection law took effect in 2020, so “none of the breached data analyzed in this paper were subject to regulation by a data protection law.” Breach notification law is the exception — California’s ran from 1 July 2003 — but notification governs disclosure after the fact, not defence. The United States still has no federal data protection statute; the ADPPA was a bill.
So the paper is evidence about what gets breached, and evidence that this dataset cannot answer what the spoke wants to know. See zero-trust and system-hardening for the control side the authors recommend (they reach for the NIST Cybersecurity Framework), and coordinated-vulnerability-disclosure for the disclosure side.
What it declines to claim
The stock-market section is honest about its own limits: it shows price trajectories for Citigroup, JPMorgan Chase and LinkedIn around their breaches and then refuses the causal reading, listing financial crises, regulatory change and macro conditions as competing explanations. A paper that publishes a suggestive chart and then argues against over-reading it is doing something the vendor material in this spoke does not.