Skip to content
Rush Commerce
AI & Automation3 min read

Google WikiSkill: an agent failure wiki lifts small models

Google Research's WikiSkill gives AI agents a persistent wiki of past failures and fixes. Smaller models with it rivaled bigger models without it. Copy the pattern.

Google Research's WikiSkill gives an AI agent something most production agents do not have: a record of what went wrong last time. The paper, by Liyan Tang and five co-authors, went up on arXiv in late August, and VentureBeat brought it to a wider audience this week. The headline result is the one an operator should care about: smaller models running WikiSkill matched larger models running without it. Better notes beat a bigger model.

What actually happened

WikiSkill splits an agent's working memory into three layers:

  • Raw layer: every execution trace, kept unchanged. Tool calls, reasoning, outputs.
  • Wiki layer: pages distilled from those traces. Failure modes, strategies that worked, a log of which skill changes helped and which did not.
  • Skill layer: the actual instructions the agent runs with, in the same file-based SKILL.md format many agent tools now use.

After each round, a maintainer agent updates the wiki, a proposer agent writes or edits a skill, and the change has to pass validation before it goes live. Skills that make things worse can be rolled back. The key design choice, per Google researcher Liyan Tang in VentureBeat: most self-improving setups throw away the diagnosis after a patch, including which fixes failed. WikiSkill keeps it.

Results across five benchmarks, as reported by The Decoder and VentureBeat:

  • Gemini 3.5 Flash went from a 49.5% average to 68.1%.
  • Qwen 27B went from 39.4% to 63.3%.
  • Gains grew with model size, but smaller models with evolved skills closed much of the gap to larger ones without them.
  • Skills evolved on a 27B model transferred to a 9B model, though transfer varied and needs checking case by case.
  • Document QA over long files barely improved.

The paper does not say whether code will be released.

Why it matters for your business

Your agent repeats mistakes because nobody writes them down. Most small-business agents run on a prompt someone tuned once. When the agent misreads a return policy or picks the wrong tool, a person fixes the output and the lesson disappears. WikiSkill is a formal version of a habit any team can start this week: keep a failure log next to the prompt.

You do not need the paper's code to use the idea. Store traces. Once a week, review failures and write each pattern as one short page: what happened, why, and the rule that fixes it. Fold the rules into the agent's skill file. Test each change against last month's failures before it ships. Keep the old version so you can roll back.

This is a cost lever. If a cheaper model with good skills matches an expensive model with none, your notes are worth real money every month.

Key takeaways

  • WikiSkill keeps raw traces, a wiki of failures and fixes, and versioned skills
  • Gemini 3.5 Flash rose from 49.5% to 68.1% average across five benchmarks
  • Smaller models with evolved skills rivaled larger models without them
  • Skill changes are validated before use and can be rolled back
  • Start a failure log beside your agent's prompt, then test every fix

Your agent keeps making the same mistake? We build agents with trace logging, versioned skills and a test set from your real failures, so fixes stick and smaller models do the work. See how we build or estimate what a cheaper model saves.

Sources: arXiv: WikiSkill, VentureBeat, The Decoder.

  • #ai-agents
  • #agent-skills
  • #google-research
  • #agent-memory
  • #small-models
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.