Key Takeaways
- OpenAI allegedly concealed evidence in a copyright lawsuit.
- The New York Times and The Daily News are challenging OpenAI’s claims.
- OpenAI’s internal searches may contradict its public statements.
- Legal actions could lead to consequences for OpenAI.
Allegations of Concealment
The New York Times and The Daily News have accused OpenAI of misrepresenting its ability to search customer chat logs and training datasets related to copyrighted material. This accusation is part of an ongoing lawsuit that has lasted two years, focusing on claims that OpenAI’s generative AI models were trained using the Times’ content without permission.
OpenAI’s Defense
Throughout the legal proceedings, OpenAI has maintained that it could not search its training data effectively. The company argued that retrieving and processing its extensive collection of ChatGPT conversations would be technically challenging and could compromise user privacy. The news outlets sought access to this data to assess whether their copyrighted journalism was included in OpenAI’s training set and how often ChatGPT generated responses based on their content.
Revelations from Deposition
In a deposition ordered by the court, OpenAI data privacy engineer Vinnie Monaco reportedly disclosed that the company had previously conducted internal searches of its training corpus to identify copyrighted journalism. This included a database of approximately 78 million de-identified ChatGPT conversations, which OpenAI used to evaluate potential copyright infringements.
Concerns Over Data Handling
Additionally, Monaco’s testimony indicated that OpenAI implemented a “Bloom” filter as part of a project aimed at monitoring output for potential copyright violations shortly after the lawsuit was initiated. The plaintiffs had initially requested a sample of 120 million chat logs, but OpenAI negotiated this down to 20 million. However, the submitted sample was criticized for excessive redactions, rendering it largely unusable according to the court.
Legal Actions Requested
The New York Times and The Daily News are now asking the judge to take action against OpenAI for allegedly withholding evidence and disrupting the discovery process. They seek to prevent OpenAI from using the 20 million chat log sample as evidence, asserting that it is unreliable. They also want the court to recognize that ChatGPT logs likely contained significant instances of their content and to compel OpenAI to cover legal fees incurred in pursuing this evidence.
OpenAI’s Response
In response to these allegations, OpenAI spokesperson Drew Pusateri denied any wrongdoing, claiming that the Times is attempting to invade user privacy as their case weakens. Pusateri stated, “As the Times’ case weakens and they’ve been forced to drop claims against us, they’re persisting with their efforts to invade the privacy of people who have nothing to do with this case, including by making these blatantly false allegations. We’ll continue defending our users’ privacy and the long-established principles of fair use.”
