
On June 18, 2025, Cybernews.com published a report on the discovery made by security researcher Bob Diachenko. He found an unsecured, publicly accessible database on the Internet, containing 30 different data sets, which was said to contain up to 16 billion logins to various services.
Here’s what we know:
The database was only accessible for a “short period of time”, during which it was not possible to establish who was behind its creation.
The login credentials are believed to come at least in part from so-called infostealer malware, harvested directly from infected users' devices. No apparent third-party website breach occurred.
The database was no longer accessible at the time of writing, so all information is based on a few screenshots published by Cybernews.com.
The story quickly spread across media outlets, and the “16 billion credentials leak” headline took on a life of its own.
We decided to take a closer look to help you understand what actually happened and what the risks are.
TL;DR: This is just another large compilation of old data, partly from infostealers, and full of duplicates. While massive, it doesn’t represent any significant new threat. The headlines were greatly exaggerated. That said, infostealer malware remains a serious risk, and prevention is key.
The discovered database contains 30 datasets with varying names and formats. Some have distinctive labels that allowed us to attribute them to specific sources or services, while others are more generic, named simply “leaks”, “log”, or “csv_data”.
Some of these names directly suggest a link to infostealer malware. For others, we can’t say for sure — we don’t have access to the full original content.
By comparing partially redacted screenshots from the original article with our database, we connected some of the credentials shown to info-stealer breaches published on Telegram and elsewhere several years ago.
Two of the databases shown in the screenshot are named “ctionbudget-unamepass.” The name - and the data format - suggest they were created using an open-source tool called Openhunting CTI.


These two databases contain 343,972,941 and 60,448,289 records, respectively, together accounting for over 404 million of the records included in the “leak.”

Openhunting CTI is a tool that observes a predefined list of Telegram channels, extracts credentials from .txt files (these are called ULP lists, and we previously covered them in this blog post), and saves these in the ElasticSearch database.
The tool itself was last updated two years ago.
We’ve checked some of the credentials in the screenshot and discovered that these credentials had already been published back in 2023.


Another database, named “logins,” contains over 184 million records. According to Cybernews.com, this dataset was already covered by Wired magazine in May.
Again, we compared the credentials with our own records and confirmed that some had already been published in 2022.


The last database we were able to analyze from the screenshots is called “breach-files,” and its format appears to be the same as “ctionbudget-unamepass.”
Once again, we were able to link some of the credentials to breaches that occurred in 2023, clearly indicating this is not newly published data.


One of the databases with a distinctive name is called “dyxless”. A quick search suggests it may be related to a Russian Telegram bot of the same name, which provides intel on individuals and legal entities.

The project has its own website with basic information, however, the bot itself is not reachable by the username published there anymore.

As is often the case on Telegram, multiple similarly named accounts exist, trying to scam the user for money.

With a quick search on a darknet forum, we discovered the “correct” address of the bot and tested its functionality - hoping to link its data sources to infostealer breaches found in our database.

Translated:
🔎 Welcome to the Duhless search engine.
The service is a tool for searching for information about individuals and legal entities and uses the latest and most unique databases for searching.
The service operates in real time and generates a report “on the go”, that is, without saving all the information received from the databases.
Your request remains anonymous, since the service does not collect information about users. The databases are updated daily and make it possible to collect a full package of information on an individual or legal entity.
The technical support of the service works in a permanent mode, and allows you to solve any issue related to the operation of the service.
We wish you a pleasant use. Best regards, the Duhless team.

We tested several dozen email addresses from older infostealer logs, including those we used to trace credentials from aforementioned datasets, but mostly with no success.
In a few cases, we were able to obtain information for an email also found in an infostealer breach.
However, the data consisted mostly of old information scraped from public leaks, with no real connection to the infostealer data in our database or the other datasets that make up the “16G leak”.


Dyxless bot does not seem to utilize infostealer logs, which leads us to two conclusions:
After the original article was published and the media frenzy that followed, the issue was discussed on several darknet forums.
Some users obviously tried to obtain the database, while others reached the same conclusions we did: that it was overhyped and largely useless.
For example, here’s a statement on the matter by one of the older members of one such forum:

Translated:
The news feeds are full of headlines about the “greatest leak” and the compromise of 16 billion passwords. The original source was the Cybernews portal, the information was quickly picked up by Forbes and other major media. It sounds like a real cyber apocalypse, but if you separate the wheat from the chaff, the picture is completely different.
Let’s start with the original source. Cybernews does find vulnerabilities and leaks from time to time. Researchers found a large data set that was temporarily accessible through unprotected Elasticsearch or object storage instances. The problem lies in the presentation and subsequent media interpretation. The use of wording like “plan for mass exploitation” and “fresh, ready-to-use intelligence” creates the false impression of a one-time and catastrophic hack of a super-duper global storage. Without technical expertise, journalists took this presentation at face value. The figure of 16 billion became a trigger, and the confusion in terms, when some write about “accounts” and others about “passwords”, only increased the media hype.
What is the fundamental error in such a presentation? As sober analysts have correctly noted, this is not a leak, but a compilation. This array is the result of many years of work by malware of the infostealer class. The mechanics of the process are simple: thousands of computers are infected, and then the infostealer steals all the passwords saved in the browsers, after which the logs are aggregated in a standardized format (URL: login: password). What the researchers found is just a giant “warehouse” of such logs. It contains a huge number of duplicates, outdated passwords and long-compromised data.
Instead of preparing a low-key technical report for the information security community with clear details (the statements are serious), they packaged their discovery in a sensational wrapper for the general public and the media.
If any new inputs, details (something interesting) appear, or analysts publish convincing evidence that the “leak” deserves attention, I will make another post, but there is no point in discussing it otherwise.
a.k.a SecAtor
The word is, the database had been listed for sale since late May 2025 on xmrbazaar.com for a mere $400. We weren’t able to confirm this directly, as the listing has since been removed, but we managed to find a screenshot:
![[object Object]](/16-billion-leak/image17.webp)
The listing received some positive reviews, and members of various darknet forums also agreed that this was the source of the “leak.”
The publication date also matches.
If it was, the $400 price tag is itself an indicator of the content’s quality. If the dataset had included any fresh, previously unpublished data, the price would have been significantly higher.
Long story short, the “16G leak” was heavily overhyped and does not appear to represent any new or significant threat. Such “leaks” shouldn’t get so much publicity, because they happen all the time.
Journalists, users, and organizations concerned with cybersecurity should instead focus on the real problems the infostealer malware poses.
Confused by headlines? You're not alone. At HackedList.io, we help organizations and individuals cut through the noise.
We monitor live stealer logs, darknet marketplaces, and private breach sources – so you get real, actionable insights instead of media-driven panic.
Want to know what’s actually putting you at risk? Enter your domain below to find credentials compromised by info-stealer malware: