The Part Nobody Wants to Hear
If you have had an email address for more than a few years, it is in a breach somewhere. That is not a prediction. It is arithmetic. Have I Been Pwned, the most conservative breach index running, passed 13 billion breached accounts years ago and keeps climbing. There are only about 4.5 billion email-using humans on the planet. The average address is not in one breach. It is in several.
I say this to clients and watch them go through the same three stages every time. First, disbelief: they use a password manager now. Then, bargaining: the breach was years ago, they changed the password. Then the part that actually matters, which is realizing that none of those defenses change what the historical record shows. The data from 2016 is still data. It still shows the email you used, the password pattern you favored, the services you signed up for, and sometimes the physical address and phone number you had at the time.
Breach data is not a live feed of someone's accounts. It is a fossil record. And fossils are exactly what investigators work with.
How the Data Gets Loose
Three pipelines feed the breach data ecosystem, and they are worth understanding separately because they produce different kinds of records.
Direct breaches. A company gets compromised, its user database gets exfiltrated, and eventually that database surfaces. Sometimes it gets dumped publicly on a forum. Sometimes it gets sold privately for years before leaking. The LinkedIn breach happened in 2012 and the full dataset did not circulate widely until 2016. A four year lag between compromise and public availability is common, which means the breach you have not heard about yet may already include you.
Combo lists. This is the part people underestimate. Someone takes thousands of individual breach databases and merges them into giant credential lists. Collection #1, which surfaced in January 2019, held 773 million unique email addresses assembled from thousands of separate sources. The COMB compilation in 2021 held around 3.2 billion email and password pairs. RockYou2021 distributed 8.4 billion password entries. None of these were new hacks. They were repackaging at scale, and repackaging is why old breaches never really die. The data gets copied, cleaned, deduplicated, and redistributed until it is effectively permanent.
Infostealer logs. The newest and in some ways ugliest source. Malware families like RedLine, Raccoon, and Vidar infected millions of machines and pulled saved passwords, browser cookies, autofill data, and session tokens straight off the device. Those logs get sold in bulk on Telegram channels and darknet markets. Law enforcement has taken chunks of this infrastructure down, including the RedLine seizure in Operation Magnus in late 2024, but the logs already sold stay in circulation. Infostealer data is different from breach data because it often includes passwords in plaintext along with the exact URL each one belonged to, captured from the victim's own browser.
What a Record Actually Contains
Run an email through a breach search engine and the raw result varies wildly by source. A minimal record is just an email and a hashed password from a forum breach in 2013. A rich one from a marketing database breach can include full name, phone number, physical address, date of birth, and IP address. An infostealer record can include plaintext credentials paired with the sites they open.
The fields matter less than the metadata investigators extract from them:
Password patterns. One leaked password tells you a string. Three leaked passwords from three different breaches tell you a system. People build passwords from formulas: a base word, a year, a site-specific suffix, a favorite number. When you can see the formula across breaches, you understand how the subject thinks about credentials. That is useful for assessing exposure, not for breaking into anything.
Service footprint. The breaches an email appears in map the services that person registered for. If the address shows up in a breach of a niche forum, a gaming service, and a defunct freelance platform, you have a partial account inventory built entirely from their mistakes. Each service is a potential pivot to a username, a profile, a post history.
Timeline anchors. Breach dates tell you an email was active and in use by a certain year. That sounds trivial until you are trying to establish whether an address found in a document was real, abandoned, or never belonged to the person claiming it. An address appearing in a 2015 breach with a 2012-era password was a working account. That is a fact you can build on.
What It Does Not Tell You
This is where most people overread the data, and where bad investigations go wrong.
It does not prove current access. A leaked password from 2018 is a historical artifact. Treating it as a live credential is the breach-data version of treating a decade-old address from an aggregator as a current residence, the exact failure mode covered in how investigators read public records. Stale is stale, whether it comes from a data broker or a database dump.
It does not prove ownership of everything attached to the email. Breach records merge data from whatever source got compromised. If a marketing database wrongly associated a phone number with an email address, that wrong association is now permanently part of the breach record. You will see it repeated across every breach search engine that indexed the same compilation, which makes it look confirmed. It is not. It is copied, and the distinction matters the same way it does with confirmation bias in any investigation.
It does not prove exposure by absence. A clean result means the address is not in the collections that service indexed. Breaches that were detected but never leaked, sold privately and kept private, or disclosed without the data escaping will not appear. Some companies never disclose at all. Absence of evidence in a breach search is not evidence of absence. It is evidence you checked one layer.
How Investigators Actually Use It
In a real workflow, breach data is a pivot engine, not an answer key.
The standard play: take a subject email, run it through breach indexes, and extract what the records give you. Old usernames embedded in the records become username searches. Password patterns become exposure assessments if the subject is a client. Associated phone numbers get run through carrier and reverse lookups. Physical addresses get checked against property and voter records. Every artifact spawns a separate line of inquiry through a separate source. The breach record itself proves nothing final. It tells you where to dig next.
Where it gets powerful is corroboration. If you have independently established that a subject used a particular handle in 2019, and a breach record from that era ties the same handle to the email you are investigating, the two facts reinforce each other. Neither one alone is conclusive. Together they start to look like identity resolution. That layered confirmation is the whole discipline, and it is why I built OMERTÀ to treat every source as a lead that has to be verified before it counts as a hit.
If you are doing email-led work, the mechanics of pivoting off an address are covered in the email OSINT guide, and the tradeoffs between the lookup services themselves are in the email tool comparison. Breach data slots into that workflow as one source among many. It is loud, it is rich, and it is seductive, because a plaintext password feels like a skeleton key. It is not. It is a historical document.
The Legal Line
Read this part twice, because it is the difference between research and a felony.
Searching indexes of already-circulating breach data for identifiers you are investigating is standard OSINT practice and generally lawful. The data is public in the practical sense: it is indexed, searchable, and accessible without bypassing any security control. Courts and clients routinely accept breach-derived leads as investigative material.
Using those credentials is a different universe. Logging into an account with a leaked password is unauthorized access under the Computer Fraud and Abuse Act in the US and equivalent laws nearly everywhere else. It does not matter that the password was public. It does not matter that the login worked. Access without authorization is the crime, and "the password was right there" is not a defense. The same applies to session tokens from infostealer logs, which are arguably worse since they bypass login entirely.
The rule is simple: breach data is for reading, not for using. The moment you authenticate as someone else, you have stopped doing OSINT and started committing a crime. Every competent investigator I know treats that line as absolute, because it is.
The Point
Your passwords are already out there. So are mine, and so are your subject's. That is not a reason for despair and it is not a reason for complacency. It is the baseline condition of doing identity work in 2026.
Breach data is a fossil record of people's account habits, and it reads like one: rich in detail, frozen in time, and easy to misinterpret if you forget when it was laid down. Use it for pivots. Use it for corroboration. Use it to understand how a subject built their digital life. Never use it to open a door, and never treat a single record as proof of anything beyond its own existence.
The investigators who get this right are not the ones with access to the most data. They are the ones who remember what the data actually is.
FAQ
Is it legal to search breach data?
Searching databases that index already-public breach material for your own accounts or for research is generally lawful. Using leaked credentials to access accounts you do not own is illegal under computer crime laws in the US and most other jurisdictions, regardless of how easy the data is to find.
If my email is not in a breach search, am I safe?
No. A clean result means your address does not appear in the collections that particular service indexed. Breaches that were never disclosed, never leaked publicly, or circulate in private channels will not show up. Absence from breach data is not evidence of safety.
What does a breach record prove in an investigation?
It proves the data existed in that collection and was tied to that identifier at some point in the past. It does not prove the password is current, that the person still uses the email, or that any linked account belongs to them today. Breach data is a historical record and a lead, not a conclusion.
How big is the breach data problem really?
Have I Been Pwned passed 13 billion breached accounts years ago. The COMB compilation alone held around 3.2 billion email and password pairs in 2021, and infostealer malware keeps adding fresh records daily. The realistic assumption for any long-used email address is multiple exposures, not zero.
Search 698 platforms and 30 data sources with verified, confirmed-only results.
Open OMERTÀ