Security.io Intelligence DeskMonday, 10 August 2026
Independent analysis
for security executives
The Security.io DailyThe Monday Intelligence Edition
Free to readers
Supported by underwriters
AI Security · Executive briefing

Self-evolving agent skills create a trajectory-poisoning control gap

New controlled research shows how attacker-supported operating trajectories can be converted into persistent agent skills, challenging trust in retained experience and automated self-improvement.

AI SecurityApplication SecuritySecurity Leadership
Why it is in today’s brief

The paper was posted on August 6, just before the requested window, and is included as a developing-risk exception rather than a reported enterprise incident. Its marginal executive value is architectural: it tests the assumption that successful prior agent trajectories are trustworthy training material. The result changes design-review priorities for self-improving agents by requiring provenance, approval and rollback around learned skills, not only model and prompt controls.

Read first

The research demonstrates a control problem in agents that learn reusable skills from stored trajectories: apparently successful experience can become a poisoned instruction source.

Act now

Inventory agents that retain trajectories or generate reusable skills.

Accountable owner

CISO with the AI platform owner, application-security leader, model-risk function and engineering architecture board

Decision horizon

This week for production inventory and safeguards; before any new self-evolving agent deployment

AssessmentDeveloping assessment
Emerging riskIndependent reproduction, production observations, identified affected frameworks, model-specific results or evidence that poisoned skills cross tenant or repository boundaries.

What happened

The paper was posted on August 6, 2026. Attribution posture: The trajectory-poisoning paper reports a controlled academic experiment and does not attribute real-world activity to any actor. PoisonedEvolution was evaluated in SkillClaw using inert canary specifications. Researchers configured attacker-supported trajectories at 10% support in the controlled experiment. The self-evolving system mechanically converted retained trajectories into skills carrying the target behaviour in 546 of 600 trials. At 10% attacker support, the experiment reported a 91.0% success rate across 546 of 600 trials.

The paper’s search-visible abstract refers to six mainstream LLM evolvers but does not identify the underlying model names in the cited excerpt. The experiment concerns systems that learn or refine reusable skills from prior operating trajectories, rather than a frontier model acting without configuration. The cited paper did not publish real-world victim indicators because the evaluation used inert canary specifications. No enterprise compromise, malicious campaign or production exploitation is established by this source. The cited source did not publish the underlying model detail described as Underlying model names in the accessible excerpt.

Why this matters now

An agent’s retained experience can become part of its effective software supply even when it is not packaged as conventional code. If a learning system converts successful trajectories into reusable skills, an attacker may not need to alter the base model or application repository; influencing the experience selected for retention may be enough to shape later behaviour. Traditional prompt filtering does not by itself govern that durable state.

The executive decision applies only where agents can learn, retain or promote skills. Stateless assistants and systems without trajectory-driven updates are not shown affected by this experiment. For qualifying deployments, however, trajectory stores and generated skills require the same governance expectations as code changes: trusted provenance, review, testing, least privilege, versioning and dependable rollback.

The decision for security leaders

Require an inventory of agents that retain trajectories, generate skills, fine-tune behaviour from operational history or import third-party skills. Architecture owners should document who can submit experience, what criteria promote it, where generated artefacts are stored and which tools those artefacts can invoke. Production promotion should be blocked unless provenance and approval are recorded.

Commission a controlled validation using inert canaries rather than assuming a framework is resistant. Tests should measure whether maliciously supported trajectories influence durable skills, whether review controls expose the change and whether rollback removes derived state. Any agent with command, repository, deployment or cloud-control-plane access should require a stricter approval path and explicit exception ownership.

Evidence of closure

  • The agent inventory identifies every trajectory-retention and skill-generation feature.
  • Promotion logs show human approval and provenance for production skills.
  • Canary testing shows poisoned trajectories are rejected or quarantined.
  • Rollback testing restores a known-good skill set without residual changes.

The Security.io assessment

This is proof-of-concept research, not evidence of an active campaign. The result is nevertheless decision-relevant because it identifies a governance surface that many AI programmes do not yet treat as executable policy: accumulated experience. The reported metric should not be generalised to every agent, framework or model because the accessible evidence does not provide model-level results or production conditions.

The appropriate response is targeted design assurance, not a blanket suspension of agent deployments. Organisations should first establish whether self-evolution exists in their architecture. Where it does, closure requires a testable chain from trajectory provenance through skill generation, approval, deployment and rollback. If that chain cannot be reconstructed, the agent should not receive privileged tools or production command execution.

Questions for the morning meeting

  • Which production agents can convert experience into durable instructions?
  • Who approves generated skills before they receive tool access?
  • Can the organisation reconstruct every input that shaped a skill?
  • Does rollback remove both the skill and derived behavioural state?

Related intelligence

Shared decision context

Appointments, dinners & sponsored intelligence

Current paid placements · clearly separated
Registration open
Sponsor's Notice · Information Security Network

Security.io Executive Roundtable: The 2027 CISO Agenda

CISO Roundtables & Executive events

View roundtables →
Invitation only
Sponsor's Notice · NoBrowser

Security.io CISO Dinner: The Secure Browser Decision

Virtual PC's & Secure Browsers in the Cloud

Request an invitation →
Black Hat week
Paid Placement · HackerFX

Security.io at Black Hat: Daily Intelligence Briefing

Catch the Daily News Where it Happens First

Follow the Black Hat desk →