OpenAI has disrupted a coordinated campaign to extract protected internal reasoning from its AI models on a large scale. The company took action against activity linked to more than 15,000 user accounts.
The company described the operation as an adversarial model-distillation effort, in which attackers systematically sought model outputs or reasoning traces that could help reproduce, train, or enhance another AI system.
The campaign’s earliest activities were observed during the first week of July 2026 and initially had a low volume. However, on July 24 and 25, there were significant spikes, with attackers sending approximately 16,000 requests using a specific extraction pattern across more than 4,000 users.
OpenAI emphasized that these figures represent attempts at reasoning extraction rather than necessarily successful ones. Following an expanded investigation, OpenAI identified related prompt behavior across a broader group exceeding 15,000 accounts and fully disrupted the network by July 28.
Unlike a conventional security breach, this activity did not involve breaking encryption, accessing an internal database, or compromising stored customer conversations.
Instead, the operators allegedly manipulated model interactions to reveal hidden reasoning artifacts in a way that was visible to requesters. One technique observed involved copying encrypted reasoning from one conversation and submitting it in another, along with instructions to decrypt and transcribe the concealed content.
Protected reasoning refers to the internal record a model uses to process a task before generating a user-facing answer. Exposing it can reveal information intentionally left out of the final response.
It could help an adversary emulate a model’s advanced capabilities. OpenAI stated that this extraction activity violated its terms of service and is not a security issue exclusive to its platform. The company warned that similar techniques could potentially affect other frontier AI systems.
Independent security researchers have also reported related attack paths through responsible disclosure. OpenAI validated these findings and noted that the research helped them understand the broader class of attacks and accelerate mitigation efforts.
This incident highlights a growing AI security concern: attacker-controlled prompts and reusable context can turn model-to-model interactions into a channel for extracting unintended information.
OpenAI attributed a significant portion of the activity to individuals associated with Moonshot AI, the developer of Kimi. However, the company clarified that it could not definitively determine whether all observed operators during this period belonged to a single actor. The attribution stops short of alleging that Moonshot AI directed the entire campaign.
Regarding mitigations and impact, OpenAI reported that it blocked or restricted fraudulent accounts, strengthened signup and infrastructure controls, and expanded monitoring for related account networks.
It also closed a replay pathway through which someone possessing another user’s encrypted reasoning could access its contents, while adding checks to detect and hold streamed output that could potentially expose hidden reasoning.
This incident underscores the implications of adversarial distillation beyond intellectual property concerns. Reasoning extracted from a protected model could transfer advanced capabilities without inheriting the original developer’s safety controls, especially in dual-use contexts.
OpenAI said it has shared relevant intelligence through the Frontier Model Forum and government information-sharing channels, and it will continue to strengthen tool defenses, classifier coverage, model refusals, and protections for partner-hosted deployments.
Cut every SOC alert investigation by 21 min. Power your SOC with instant IOC context for immediate response: Integrate TI Lookup in your SOC
Toronto, Canada, October 8th, 2026, CyberNewswire Insignary Launches Clarity AIR: Closing the Blind Spot Between…
A threat actor published a malicious version of the tensorlake npm package on October 8,…
A proof-of-concept (PoC) exploit has been released for CVE-2026-102489, a critical vulnerability in Zammad that…
A critical vulnerability in LMCache allows unauthenticated attackers to execute arbitrary code against reachable multi-process…
16 malicious Firefox extensions that impersonate cryptocurrency wallets to intercept recovery phrases and private keys…
Exposed directories on five servers have revealed an operational DarkSword/Coruna exploitation platform built to compromise…