Lost amid the news cycles on OpenAI's disclosure about poorly contained AI models that went on to hack into HuggingFace and other companies was this disclosure from the German health insurer Universa, which said OpenAI scraped customer data while it was supposedly unprotected due to a misconfiguration during an IT migration.
https://www.heise.de/en/news/uniVersa-OpenAI-AI-crawler-accessed-customer-data-11375869.html
I reached out to Universa to learn if they knew how many records were accessed, but they ignored that question and sent this statement:
"We can confirm that there was an IT security incident at uniVersa. During an IT migration, a security setting was accidentally omitted temporarily on one of our IT systems. As a result, access to data on a server was possible for a few hours. Once the incident was detected, the unwanted access to the affected server was shut down immediately. The relevant data protection and supervisory authorities were informed without delay. During this timeframe, data from this server was retrieved by an automated web crawler belonging to a US artificial intelligence provider (OpenAI, L.L.C.). Therefore, this was not a targeted attack by criminals, but rather an automated process that regularly occurs on the internet.
Following our immediate contact, OpenAI has already confirmed that the data retrieved by its automated web crawler was not used for training its AI models and is ensuring that this will not happen in the future either.
The incident affected a single server intended for automated data exchange with sales partners. General personal data, such as names and addresses, as well as contract data, such as insurance numbers and policy information and, in the case of some customers, bank details (IBAN and BIC), were stored there. Particularly sensitive information, such as health, login, or credit card data, as well as the central administrative and data systems and the customer portal, were not affected.
In coordination with the relevant authorities and IT security specialists, we immediately began contacting those affected in writing to inform them about the most important aspects of the incident for them. This was our top priority. To protect the individuals affected, we are not providing any further details at this time.
The misconfiguration was rectified immediately, our security measures were reviewed, and further ad-hoc measures were taken. The external IT forensic investigation into the incident has been completed. Based on the forensic findings, we will be implementing additional technical and organizational measures."
To me, this incident raises a question that is far more scary than some "rogue" AI agent creating work-once malware that doesn't even try to hide (btw, read this from Charlie Eriksen, who thinks he found the malware OpenAI agents wrote: https://www.aikido.dev/blog/anthropic-rogue-agents-package-stole-keys)
The question is, how many times a day or week do agents working for on behalf of OpenAI, Anthropic, xAI, CoPilot, etc. ingest sensitive and protected financial and health data? My guess is quite a lot.