# Agent Skills Evidence Base: Claims Register

Every quantitative claim and direct quotation behind the Agent Skills research pages, with the paper and page it came from, grouped by source so a single paper can be checked in one pass.

**What this is.** 299 claims pulled from a cross-paper analysis of the Agent Skills literature. 195 of them were load-bearing enough to re-derive from the source PDFs across five independent adversarial passes, each asking what correction a hostile reader would force. 94 came back publishable, 93 publishable only with a caveat attached, and 8 did not survive and appear nowhere in the published pages.

**How to check a claim.** Open the paper at its arXiv link, go to the PDF page named in the row, and the figure should be there. Page numbers refer to PDF pages rather than printed page numbers where the two differ. Claim IDs beginning with B come from the cross-paper comparison, and those beginning with C come from the point of view built on it.

**Caveats are not decoration.** Roughly half of these figures are correct only with a specific qualifier attached, usually a sample size, a single-run status, or a conditional cohort. The published pages carry that qualifier in the same sentence as the figure, and a figure quoted without it will usually overstate what the paper found.

**If you find an error.** Tell us. We will correct it in the open and log it in the changelog on the hub page. Eight claims failed this process before publication and more will.

**Corpus.** 35 papers read in full, drawn from roughly fifty on the SKILL.md artifact posted to arXiv between 25 July and 18 August 2026, out of a 170-candidate sweep.

---


## Demystifying Agent Skills: Why They Work-Until They Don't

arXiv [2608.14036](https://arxiv.org/abs/2608.14036) · 28 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B009 | p.20 | Demystifying Agent Skills reports Skill vs Raw = +2.84 pts with a 95% CI of [-2.27, +7.95], which spans zero. |  |
| B010 | p.20 | In Demystifying Agent Skills only Skill vs Workflow Memory (+6.06, CI [+0.76, +11.36]) clears zero. |  |
| B028 | p.8 | Demystifying open-coded 240 trajectories into 12 canonical modes (kappa = 0.952, 714 human checks) and found procedural_anchor = 65.7% of skill mechanisms versus knowledge_injection = 4.5%. |  |
| B029 | p.8 | Demystifying states that skills usually do not work by supplying missing facts but by stabilizing action. | "Skills usually do not work by supplying missing facts. They work by stabilizing action" |
| B030 | p.23 | Demystifying shows skills remove environment_infrastructure_failure 5.3% -> 0.2%, background_service_lifecycle_failure 2.7% -> 0.8%, shell_code_corruption 1.1% -> 0.2% and output_format_schema_mismatch 7.4% -> 3.2%. |  |
| B031 | p.23 | Demystifying shows skills do not remove algorithmic_logic_error (8.3% -> 7.4%) or static_verification_without_runtime (12.5% -> 11.7%). |  |
| B039 | p.9, p.23 | Demystifying found skill_guidance_misapplied_or_ignored at 10.0% of skill-arm cases versus 0.8% raw. |  |
| B044 | p.10 | In Demystifying, growing pools from k = 5 to 100 drops execution-time actual-use precision from 29.6% to 3.3%. |  |
| B045 | p.10 | In Demystifying, downstream success barely moves from 36.4% to 39.3% as pools grow from k = 5 to 100. |  |
| B046 | p.10, p.11 | Demystifying isolates confusability: similar-distractor pools drop top-1 from 70.5% to 53.4%, random pools from 97.7% to 84.1% and dissimilar pools from 96.6% to 93.2%. |  |
| B052 | p.10, p.11 | Demystifying's controlled contrast shows a 100-skill pool of similar skills hurts more than a 100-skill pool of random ones, measured on offline top-1 identification precision only. |  |
| B053 | p.1 | Demystifying's abstract states that confusable distractors impair offline identification, yet downstream success remains stable. | "confusable distractors impair offline identification, yet downstream success remains stable" |
| B054 | p.10 | In Demystifying, offline embedding top-1 precision degrades gently from 88.3% to 76.9% while execution-time precision collapses from 29.6% to 3.3%. |  |
| B132 | p.23 | In Demystifying's matched 83-task cost table, Skill beats Workflow Memory by +4.8 pp for +95.3K tokens. |  |
| B133 | p.23 | In Demystifying's matched 83-task cost table, Skill beats Raw by +5.5 pp for -34.2K tokens. |  |
| B174 |  | Demystifying's confidence interval spanning zero sits in direct tension with SkillCorpus's positive result. |  |
| C001 | p.7, p.8 | Demystifying open-coded 240 trajectories yielding 238 retained labels, with human/LLM taxonomy-aggregation agreement of 95.8% exact and Cohen's kappa = 0.952. |  |
| C002 | p.7, p.8 | Procedural anchoring accounts for 65.7% of how skills work and knowledge injection accounts for 4.5%. |  |
| C006 | p.23 | What skills move: environment_infrastructure_failure 5.3% -> 0.2%, background_service_lifecycle_failure 2.7% -> 0.8%, shell_code_corruption 1.1% -> 0.2%, output_format_schema_mismatch 7.4% -> 3.2%. |  |
| C007 | p.23 | What skills do not move: algorithmic_logic_error 8.3% -> 7.4% and static_verification_without_runtime 12.5% -> 11.7%. |  |
| C022 | p.10 | As pools grow from 5 to 100 skills, execution-time actual-use precision falls from 29.6% to 3.3%. |  |
| C023 | p.10-11 | At k = 100, top-1 precision on similar-distractor pools is 53.4%, on random pools 84.1% and on dissimilar pools 93.2%. |  |
| C024 | p.1, p.10 | In the same controlled sweep downstream success did not follow the precision collapse, moving 36.4% -> 39.3%. |  |
| C025 | p.1, p.10 | The authors state that exact ground-truth skill invocation is neither sufficient nor strictly necessary for success. | "exact ground-truth skill invocation is neither sufficient nor strictly necessary for success" |
| C034 | p.9 | Demystifying finds skill_guidance_misapplied_or_ignored at 10.0% versus 0.8% without skills. |  |
| C053 | p.23 | Skills beat workflow memory by +4.8 pp for +95.3K tokens. |  |
| C098 | p.10-11 | Similar-pool top-1 precision is 53.4% versus dissimilar-pool 93.2% at k = 100. |  |
| C099 | p.5, p.28 | Failure-only pools drove skill performance below the raw baseline, and withholding outcome labels collapsed mixed pools from 0.7462 to 0.4000 in one cell. |  |

## SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents

arXiv [2607.15557](https://arxiv.org/abs/2607.15557) · 22 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B001 | p.6 | SkillCorpus reports +7.5 +/- 2.3 pp on SkillsBench (z = 3.2), +1.51 on GDPVal and +2.79 on QwenClawBench, pooled over 72 main-grid runs (74 including the single-run frontier check), across 4 harness x backbone cells. |  |
| B002 | p.6 | SkillCorpus's frontier check moved Claude Opus 4.7 from 39.1% to 47.1% (+8.0 pp) in a single run. |  |
| B019 | p.5 | SkillCorpus is explicit that effect size tracks headroom: SkillsBench starts near 10% and gains +7.5 pp, while GDPVal and QwenClawBench start at 65-85% and gain +1.51 and +2.79. |  |
| B021 | p.7 | SkillCorpus bins SkillsBench tasks by top reranker score and finds mean delta of +2.2 pp, +6.2 pp and +25.1 pp across three match bins, surviving control for task difficulty (partial r = 0.34-0.40). |  |
| B022 | p.7 | SkillCorpus finds that where coverage is thin the gain floors at zero rather than going negative, which it calls a supply problem rather than a retrieval one. | "a supply problem rather than a retrieval one" |
| B023 | p.7 | The same SkillCorpus corpus and selections gave +13.4 pp on Raven and +5.8 pp on OpenClaw with an identical backbone. |  |
| B024 | p.7 | SkillCorpus's qualitative trace reading, which the authors explicitly flag as not an established mechanism, is that one harness completes an execute-verify-fix loop while the other stops after writing scripts it never runs. | "not an established mechanism" |
| B025 | p.7 | SkillCorpus states that skill utility appears to be a joint property of the corpus and the harness. | "Skill utility appears to be a joint property of the corpus and the harness." |
| B034 | p.12 | SkillCorpus, from two worked case studies rather than a full-corpus analysis, states that having a recipe displaced the improvisation that succeeds without one. | "having a recipe displaced the improvisation that succeeds without one" |
| B035 | p.13 | SkillCorpus summarises its mechanism finding as skills transfer procedures, not topics. | "Skills transfer procedures, not topics" |
| B043 | p.12 | SkillCorpus's one documented regression case states that the skills are on-topic, but their procedures do not extend to this task's structure. | "The skills are on-topic, but their procedures do not extend to this task's structure" |
| B134 | p.13 | SkillCorpus's serving overhead on its 87-task Opus 4.7 frontier check, estimated from released eval-cache token counters, is 158K tokens/task with skills versus 122K without, about 30%. |  |
| B135 | p.13 | SkillCorpus's ingest cost is about 269,000 LLM-judge calls for one curation pass. |  |
| B144 | p.3 | SkillCorpus independently measured 59.7% exact duplication at its dedup stage plus 13,268 confirmed semantic near-duplicates. |  |
| B150 | p.10 | SkillCorpus states that the composite score does not predict per-task pass rate on any benchmark (all \|r\| < 0.10, p > 0.35). | "The composite score does not predict per-task pass rate on any benchmark (all \|r\| < 0.10, p > 0.35)" |
| B173 |  | SkillCorpus's +7.5 pp result sits in direct tension with papers reporting nulls, and the resolution offered is that any unconditional claim is unsupported once headroom, coverage-match and harness are conditioned on. |  |
| B179 |  | SkillCorpus reports quality scores not predicting utility, with \|r\| < 0.10. |  |
| B187 |  | SkillCorpus finds Opus 4.7 still gains +8.0 pp, contradicting the view that frontier models make skills redundant. |  |
| B189 | p.8 | SkillCorpus explicitly does not study temporal churn and calls its release a 2026-Q2 snapshot. | "a 2026-Q2 snapshot" |
| C003 | p.13 | SkillCorpus states that skills transfer procedures, not topics. | "Skills transfer procedures, not topics" |
| C054 | p.13 | Serving overhead is about 30%, at 158K versus 122K tokens per task. |  |
| C084 | p.10 | An unvalidated composite quality score gives \|r\| < 0.10 on every benchmark. |  |

## Ratchet: How Reliable Must an LLM Judge Be to Retire a Skill?

arXiv [2605.22148](https://arxiv.org/abs/2605.22148) · 18 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B006 | p.9 | Ratchet reports a rolling gain of +0.328 +/- 0.018 over 100 rounds on MBPP+ hard-100 with 3 seeds. |  |
| B026 | p.10 | Ratchet's travel study gives within-family gains of +0.326 (Opus 4.7), +0.168 (Kimi), +0.105 (GLM-5), +0.087 (DeepSeek V3.2), +0.020 (Qwen3-Coder), -0.010 (Mistral Large 3) and -0.012 (GPT-5.5). |  |
| B087 | p.10 | Ratchet reports the A4-style ablation lowering N_min from 100 to 20 with a zero retirement threshold yields -0.019 +/- 0.010, below the no-skill baseline across all three seeds. |  |
| B088 | p.3, p.8 | At N_min = 20 the Hoeffding deviation is epsilon approximately 0.44, so a genuinely useful skill can be retired on unlucky draws. |  |
| B089 | p.3, p.8 | Ratchet models the grader as a binary channel with false-pass rate rho_F->P and phantom-failure rate rho_P->F, requiring for eviction that true pass rate p-bar(s) <= (pi_tau - rho_F->P)/kappa where pi_tau = (1-tau)/2. |  |
| B090 | p.3, p.8 | Ratchet shows nothing is evictable at any sample size once rho_F->P >= (1-tau)/2, which is 0.45 at tau = 0.10, and everything is evictable once rho_P->F >= 1 - pi_tau = 0.55. |  |
| B091 | p.8 | In Ratchet, phantom failures degrade the guarantee smoothly through a divisor while false passes disable it additively and discontinuously. |  |
| B092 | p.8 | Ratchet states that a judge with a 40% error rate all in the harmless direction is safer than one with 46% in the harmful one. | "A judge with a 40% error rate all in the harmless direction is safer than one with 46% in the harmful one" |
| B093 | p.8 | In Ratchet, epsilon and rho-bar_F share a numerator and only epsilon shrinks with more samples. |  |
| B099 | p.11 | Ratchet reports the audited real judge at rho_F->P approximately 0.01, rho_P->F approximately 0.95 and kappa approximately 0.04. |  |
| B192 | p.6 | Ratchet notes no task ever receives two skills and its curator cannot detect near-duplicates except through outcomes. |  |
| B197 |  | Ratchet flags as explicitly out of scope the regime where a library learns phrasing that induces judge blindness so the false-pass rate climbs along its own trajectory. |  |
| C008 | p.10 | Ratchet's within-family gains run from +0.326 for Opus 4.7, on a task slice filtered by Opus's own failures, down to -0.012 for GPT-5.5, which the authors attribute to a baseline already elevated at 0.212. |  |
| C009 | p.10 | Ratchet reports -0.010 for Mistral Large 3, the weak-end case where the tasks are simply out of reach. |  |
| C011 | p.10 | The model-strength result underlying the capability band is one 100-task slice. |  |
| C041 | p.3, p.8 | Once the grader's false-pass rate reaches (1-tau)/2, which is 0.45 at tau = 0.10, nothing is evictable at any sample size. |  |
| C042 | p.3, p.8 | Direction beats rate: a judge with a 40% error rate all in the harmless direction is safer than one with 46% in the harmful one. | "a judge with a 40% error rate all in the harmless direction is safer than one with 46% in the harmful one" |
| C043 | p.3, p.8 | Epsilon shrinks with samples while the bias term does not. |  |

## Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries

arXiv [2605.19576](https://arxiv.org/abs/2605.19576) · 17 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B005 | p.5 | Library Drift reports a rolling gain of +0.328 +/- 0.018 over 100 rounds on MBPP+ hard-100 with 3 seeds. |  |
| B081 | p.3 | Library Drift defines drift as E[pass@1 \| S_t] < E[p0] for some t > 0, a property of the set rather than any member, so drift can occur even when every individual skill appears reasonable. | "drift can occur even when every individual skill appears reasonable" |
| B082 | p.3 | Library Drift names three sub-modes: stagnation, bloat and erosion. |  |
| B083 | p.1 | Library Drift states that end-task metrics decline gradually with no explicit error signal. | "End-task metrics decline gradually with no explicit error signal" |
| B084 | p.4-5 | In Library Drift's induced-erosion condition the bank collapses to 2 active skills by round 30 while aggregate pass@1 declines only gradually, because the router adaptively selects NONE. |  |
| B085 | p.4-5 | In Library Drift, router engagement drops to 18.9% in the same window, so the trace-level signal is available before the score-level one, though the paper reports no crossover point and quantifies no lead time. |  |
| B086 | p.3, p.5 | Library Drift's ablation A4, lowering the evidence requirement from N_min = 100 to 20 and tightening the retirement threshold to zero, yields -0.019 +/- 0.010, below the no-skill baseline, consistent across all three seeds. |  |
| B139 | p.5 | Library Drift reports 14.5k LLM calls per 100 rounds, 43% above baseline, with 6.5h versus 2.3h wall time. |  |
| B183 |  | Library Drift's whole thesis of aggressive retirement is in tension with its own A4 ablation at -0.019, and at N_min = 20 epsilon approximately 0.44 makes eviction a coin flip. |  |
| B185 |  | Library Drift's A5/A6 ablations exceed the deduped default, and the authors say the effect is scale-dependent: at 50 skills a stylistic prior substitutes for dedup, while at hundreds of patterns they expect dedup to become load-bearing. |  |
| C039 | p.3 | Library drift is defined as E[pass@1 \| S_t] < E[p0], a property of the set rather than any member, so drift can occur even when every individual skill appears reasonable. | "drift can occur even when every individual skill appears reasonable" |
| C040 | p.4-5 | In the induced-erosion run the bank collapsed to 2 active skills by round 30 while aggregate scores declined only gradually. |  |
| C045 | p.5 | Lowering the evidence floor from N_min = 100 to 20 while tightening the threshold produced -0.019, below the no-skill baseline, on all three seeds. |  |
| C046 | p.5 | At N_min = 20 the Hoeffding deviation is epsilon approximately 0.44, making eviction a coin flip. |  |
| C049 | p.4-5 | Router engagement collapsed to 18.9% by round 30 in the induced-erosion run while aggregate pass@1 declined only gradually, and the paper reports no crossover point, with healthy engagement at 70-80% and below 50% a halt signal. |  |
| C088 | p.7 | Library Drift states that the bottleneck in self-evolving skill libraries is not the author; it is the librarian. | "The bottleneck in self-evolving skill libraries is not the author; it is the librarian" |
| C101 | p.5 | Router engagement below 50% means stop. |  |

## Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability

arXiv [2607.26637](https://arxiv.org/abs/2607.26637) · 15 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B058 | p.1, p.6 | Filesystem-Based Memory deliberately puts declarative content and procedural skills in one store and measures both. |  |
| B059 | p.9, p.13 | On PersonaMem 32k the agent-curated store scores 37.5% against a verbatim chronological dump's 78.1% on identical questions. |  |
| B060 | p.51 | For all 13 questions the dump answers and curation misses, the gold-critical fact is present in the curated store, with grounding 0.94-0.98 and attribution 0.97-1.0. |  |
| B061 | p.51-52 | Damage is graded by distance from the chronological record: Verbatim 78.1 > Foldered 62.5 > Agent-curated 37.5. |  |
| B062 | p.52 | Rebuilding the store with a stronger management agent (gpt-5.4) recovers the score from 37.5 to 56.3. |  |
| B063 | p.52 | The Filesystem-Based Memory authors write that they read this as a limitation of the backbone model rather than of the curated representation itself, leaving open whether the residual gap to 78.1 is a search-side problem or an irreducible cost of the curated form. | "We read this as a limitation of the backbone model rather than of the curated representation itself" |
| B080 | p.23 | Filesystem-Based Memory frames declarative memory and skills as one filesystem memory carrying different content. | "declarative memory and skills are one filesystem memory carrying different content" |
| B177 |  | Filesystem Memory reports 78.1 -> 37.5 from curation, supporting the position that curating episodic state is lossy by construction. |  |
| B190 | p.20 | Filesystem Memory states that within 140 tasks and one conversation length nothing here measures months-long accumulation. | "within 140 tasks and one conversation length nothing here measures months-long accumulation" |
| C012 | p.9, p.13, p.51 | On PersonaMem 32k the agent-curated store scored 37.5% against a raw chronological dump's 78.1% on identical questions. |  |
| C013 | p.9, p.13, p.51 | For all 13 questions the dump answered and curation missed, the gold-critical fact was present in the curated store, with grounding 0.94-0.98 and attribution near 1.0. |  |
| C014 | p.51-52 | The damage is monotone in distance from the chronological record: Verbatim 78.1 > Foldered 62.5 > Agent-curated 37.5. |  |
| C015 | p.52 | Rebuilding the store with a stronger management agent recovers 37.5 to 56.3, which the authors read as a limitation of the backbone model rather than of the curated representation itself, leaving the search-side half of the hypothesis untested. | "a limitation of the backbone model rather than of the curated representation itself" |
| C021 | p.23 | Filesystem-Based Memory argues that one substrate can serve both declarative memory and skills. |  |
| C090 |  | A stronger management agent already recovers the curated store from 37.5 to 56.3, and the open test is the search side, which the authors did not run. |  |

## Rethinking Self-Evolving Agent Skills: Feedback Dynamics over Multiple Rounds

arXiv [2608.02636](https://arxiv.org/abs/2608.02636) · 14 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B018 | p.5 | Feedback Dynamics finds only 55 of 388 candidate skill revisions (14.2%) establish a validation best. |  |
| B101 | p.6, p.17 | Across 42 controlled runs, Feedback Dynamics finds per-round new-best yield falls 26.2%, 26.2%, 19.0%, 19.5%, 4.9%, 12.8%, 5.3%, 11.1%, 8.8%, 3.0%. |  |
| B102 | p.6, p.17 | In Feedback Dynamics, rounds 1-4 produce 38 of 55 new bests (69.1%) from far fewer candidates. |  |
| B103 | p.6, p.17 | In Feedback Dynamics, 6 of the 11 finally-selected skills first appear in rounds 6-9. |  |
| B104 | p.6, p.23 | In Feedback Dynamics, SpreadsheetBench peaks at round 3 and all five later candidates score lower. |  |
| B105 | p.6, p.23 | In Feedback Dynamics, DocVQA never exceeds its parent across 23 candidates. |  |
| B106 | p.6, p.23 | In Feedback Dynamics, a DeepSeek LiveMath run drops validation from 40.0 to 11.4. |  |
| B107 | p.5-6, p.18 | In Feedback Dynamics, GPT-5.5 on LiveMath shows the only validation improvement (51.4 -> 57.1) loses 6.6 points on released test and 16.7 on transfer. |  |
| B138 | p.27-28 | Feedback Dynamics reports 2,750 target-model calls for one selected evolution run, amortising to 6.02 calls per deployed task at n = 548. |  |
| C051 | p.6 | Per-round yield collapses after round 4, yet 6 of 11 final selections first appeared in rounds 6-9. |  |
| C052 | p.12 | A byte-identical artifact scored 71.43%-83.67% across eight evaluations. |  |
| C082 | p.12 | Reruns of identical artifacts span 12 points, with SD 3.92 over eight evaluations of the same bytes. |  |
| C083 | p.29 | Two fixed verifiers disagreed on 4.8% of verdicts, changing one benchmark's measured delta from -3.0 to 0.0. |  |
| C089 | p.5 | 191 of 210 candidate revisions changed the skill, and only 29 established a validation best. |  |

## Skill Blocks: How Should an Agent Load Its Skill? A Caching-Correct Comparison of Pre-load, On-Demand Tool-Loading, Progressive Disclosure, and Hybrid

arXiv [2608.14943](https://arxiv.org/abs/2608.14943) · 14 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B121 | p.1, p.10 | Skill Blocks is the only paper that measures loading and token economics properly, and its main finding is that there is no winner. |  |
| B122 |  | For a small skill (~2K tokens) in single-turn use, hybrid stubs are best, with SearchQA at -27.4% and pure tool-loading at +48.4% (a cost). |  |
| B123 |  | For a large skill in single-turn use, hybrid is best but deletion should be tested first: SpreadsheetBench hybrid -39.8% versus static pruning -55.5%, cheaper on 236/276 cases. |  |
| B124 |  | For a small, always-needed, multi-turn skill the mechanisms are near parity: ALFWorld -12.55% at lambda=1, shrinking to -0.6% at lambda=8. |  |
| B125 |  | For large, compressible, multi-turn skills, Skill Block / hybrid win: ScienceWorld -62.5% / -52.8% and SynthProc -73.0% / -66.6%. |  |
| B126 | p.9 | In Skill Blocks, cache reads are 74-94% of raw input. |  |
| B127 | p.9 | In Skill Blocks, SynthProc's reference arm has raw input 70,528 against full's 65,380, yet an effective input of 18,428, which is 14.8% below full. |  |
| B128 | p.10 | In Skill Blocks, progressive disclosure never leads in any of the five cells. |  |
| B181 |  | Skill Blocks finds progressive disclosure never leads in 5 of 5 cells, so measured evidence beats the architectural argument in favour of hybrid stubs. |  |
| B195 |  | Skill Blocks is one provider and one production endpoint, with the authors flagging that its transfer checks are configuration transfer, not model isolation. |  |
| C057 | p.9 | Cache reads are 74-94% of raw input. |  |
| C058 | p.9 | A naive raw-token comparison ranked arms backwards in a documented case: raw 70,528 versus 65,380, yet effective input 18,428, i.e. 14.8% below the arm it appeared more expensive than. |  |
| C061 | p.1, p.12, p.16 | An exploratory, deliberately non-content-parity static-pruning arm with six sections permanently deleted came in at -55.5%, cheaper than hybrid on 236 of 276 SpreadsheetBench cases, and the authors are explicit that it is not a content-parity substitute. | "is not a content-parity substitute" |
| C093 | p.1, p.12, p.16 | A non-content-parity static-pruning arm came in at -55.5%, cheaper than hybrid on 236/276 SpreadsheetBench cases, on one benchmark only and with the authors' own caveats attached. |  |

## The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents

arXiv [2607.07436](https://arxiv.org/abs/2607.07436) · 13 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B017 | p.7 | The Blind Curator reports an end-to-end library lift of +0.014 +/- 0.054 over 3 seeds, which is undetectable. |  |
| B027 | p.13 | The Blind Curator finds that with a weaker composer, evolution never beats the no-skill floor, concluding that skills amplify a capable composer rather than teach a weak one. | "skills amplify a capable composer rather than teach a weak one" |
| B094 | p.6 | The Blind Curator reports genuine contribution-based retirements per run of 1.3 clean, 0.7-1.0 under symmetric noise at rho = 0.1-0.4, and 0.0 / 0.3 / 0.0 under bias at q = 0.2 / 0.45 / 0.7. |  |
| B095 | p.6 | In The Blind Curator raw deprecation counts never reach zero in any condition (9.7 clean, 10.0-12.0 under noise, 7.7 / 3.7 / 3.0 under bias, 10.3 for the real judge) because cap-eviction churn fills in for the stopped curator. |  |
| B096 | p.6-7 | In The Blind Curator, outcome harm is an inverted U peaking at the cliff: delta eval versus clean is -0.021 at q = 0.2, -0.065 at q = 0.45, and recovers to +0.039 at q = 0.7. |  |
| B097 | p.6-7 | The Blind Curator states that the dangerous judge is not the blindest one but the half-blind one, which feeds the synthesizer while disarming the curator. | "The dangerous judge is not the blindest one but the half-blind one, which feeds the synthesizer while disarming the curator." |
| B098 | p.5, p.15 | The Blind Curator's audited real judge has rho_F->P approximately 0.01, rho_P->F approximately 0.95 and kappa approximately 0.04, 45x clear of the fatal edge but with almost no resolution. |  |
| B100 | p.5, p.15 | The Blind Curator states that the realistic failure mode of a strict judge is signal collapse, not reward hacking. | "The realistic failure mode of a strict judge is signal collapse, not reward hacking." |
| B198 |  | The Blind Curator flags as explicitly out of scope the regime where a library learns phrasing that induces judge blindness so the false-pass rate climbs along its own trajectory. |  |
| C010 | p.13 | With a weak composer, evolution never beats the no-skill floor at all, because skills amplify a capable composer rather than teach a weak one. | "skills amplify a capable composer rather than teach a weak one" |
| C044 | p.6-7 | Outcome harm is an inverted U peaking at the cliff: -0.021 at q = 0.2, -0.065 at q = 0.45, and +0.039 at q = 0.7, because a fully blind judge starves the synthesizer too while a half-blind one keeps feeding it while disarming the curator. |  |
| C050 | p.6 | Raw deprecation counts never reached zero in any condition, including ones where genuine contribution-based retirement was exactly 0.0, because cap-eviction churn fills in. |  |
| C100 | p.9 | Before any self-evolution, inject known defects, measure the judge's false-pass rate and compare it to (1-tau)/2. |  |

## ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?

arXiv [2608.03874](https://arxiv.org/abs/2608.03874) · 11 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B007 | p.7 | ContinualSkillBench reports +16.2% relative raw improvement in 13/15 model-domain cells, single run per cell. |  |
| B032 | p.7-8 | ContinualSkillBench finds explicit skills win on rigid-output tasks, with Healthcare Programmatic 0.250 -> 0.500 and Law exact-match 0.900 -> 0.975. |  |
| B033 | p.7-8 | ContinualSkillBench finds plain in-context learning wins on open-ended rubric scoring in all three tested domains. |  |
| B069 | p.7 | ContinualSkillBench, averaged over Law/Finance/Healthcare on GPT-5.3-Codex, scores Independent 0.466, pure ICL 0.605 and skill-maintaining Sequential 0.602. |  |
| B070 | p.7 | ContinualSkillBench's RAG-over-trajectories arm reproduces the in-context-learning pattern. |  |
| B071 | p.2 | ContinualSkillBench concludes that much of the improvement can arise from adaptation to prior context and feedback rather than reusable skill abstraction alone. | "can arise from adaptation to prior context and feedback rather than reusable skill abstraction alone" |
| B117 | p.8, p.18 | GPT-4o ended ContinualSkillBench with 384 generated skills against GPT-5.3-Codex's 205, while invoking them less often in later tasks, with LLM-judged skill quality of 5.68 versus 7.94. |  |
| B191 |  | ContinualSkillBench's single-domain RAG arm is the closest thing in the corpus to a skills-versus-memory comparison, and it is one cell. |  |
| C004 | p.7-8 | ContinualSkillBench measured the boundary directly: explicit skills win on rigid-output tasks, with Healthcare Programmatic 0.250 -> 0.500. |  |
| C005 | p.7-8 | ContinualSkillBench found plain in-context learning wins on open-ended rubric-scored tasks in all three tested domains. |  |
| C018 | p.7 | Skill maintenance ties plain in-context learning, 0.602 versus 0.605. |  |

## On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification

arXiv [2608.18066](https://arxiv.org/abs/2608.18066) · 11 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B011 | p.4-5 | Fragility of Self-Improving Agents finds AWM is net negative on 2 of 3 benchmarks against a strong GPT-5-mini baseline. |  |
| B012 | p.4-5 | Fragility of Self-Improving Agents finds RBank's +1.5% has p = 0.23. |  |
| B020 | p.4 | Fragility's nulls come from deliberately starting with a stronger baseline than the original papers used. |  |
| B108 | p.4-5 | Fragility re-ran two memory-based self-improvement methods with 3 runs and 3 orders, and variance increased in 17 of 24 cases (about 71%), with best-worst gaps blowing out to 10.42%. |  |
| B109 | p.5 | Fragility finds the default benchmark order is a hidden easy-to-hard curriculum, with baseline moving-average pass rate falling from about 75% to under 40% past task ID 150. |  |
| B110 | p.1 | In Fragility, shuffling task order turns the reported +1.5% into -4.5%. |  |
| B111 | p.8-9 | In Fragility, adding rubrics, environment feedback and prompt fixes recovers 2.9 points, 31% of the degradation, leaving 69% unexplained. |  |
| B175 |  | Fragility's net-negative result sits in direct tension with SkillCorpus's positive result. |  |
| B186 |  | Fragility finds a strong baseline erases skill gains, supporting the view that frontier models make skills redundant where the bottleneck is reasoning. |  |
| C047 | p.1, p.4-5 | Variance rose in 17 of 24 self-improvement cases, the default benchmark ordering turned out to be a hidden easy-to-hard curriculum, and shuffling it turned a reported +1.5% into -4.5%. |  |
| C048 | p.1, p.4-5 | Better specification recovered only 31% of the loss. |  |

## Do Personalized Skills Help Coding Agents? An Empirical Study of Developer Interaction Histories

arXiv [2608.10319](https://arxiv.org/abs/2608.10319) · 11 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B015 | p.6 | Do Personalized Skills Help? finds personalised skills score +0.97, p = .399. |  |
| B016 | p.6 | Do Personalized Skills Help? finds a random other developer's skill scores +0.92, p = .451. |  |
| B036 | p.8 | Personalised skills produced almost no score gain but raised successful-validation runs from 43.1% to 58.9% and test command groups from 0.56 to 0.97. |  |
| B072 | p.6-7 | Personalised skills mined from a developer's own history scored -6.33 versus no skill with 0 relevant prior sessions, 0.00 with 1-2 sessions, +0.10 with 3-5 and +10.17 with 6 or more. |  |
| B073 | p.6-7 | Below the evidence threshold a pooled generic skill beat the personalised skill in every bin. |  |
| B074 | p.11-12 | Only 6 of 13 developers improved with their own skill, while the generic skill helped 11 of 13. |  |
| B136 | p.8 | Personalised skills raised agent tokens from 442,096 to 597,120 (generic 643,578), tool calls from 8.47 to 9.46 and time from 91.76 to 106.16 s. |  |
| B196 | p.5 | The Personalized Skills paper has GPT-5.5 generating the skill, playing the developer, executing the task, and judging the result. |  |
| C019 | p.6-7 | Against a generic pooled skill, the personalised skill scores -8.00 with 0 relevant prior sessions, -3.81 with 1-2, -7.20 with 3-5 and only +5.67 at 6 or more. |  |
| C020 | p.6-7 | Against no skill the same bins run -6.33 / 0.00 / +0.10 / +10.17 on bin sizes of 3/16/11/12 tasks, with the overall personalisation effect not significant (p = .399). |  |
| C055 | p.6, p.8 | Personalised skills raised agent tokens from 442K to 597K and execution time from 91.76 to 106.16 s while adding +0.97 points of score. |  |

## Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents

arXiv [2608.11888](https://arxiv.org/abs/2608.11888) · 10 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B037 | p.5-6 | Agent Skills Can Be Harmful triaged 125 confirmed functional failures and found Applicability Mismatch, the obviously-wrong-skill case, is 2 of 125 (1.6%). |  |
| B038 | p.5-6 | Agent Skills Can Be Harmful found Task-Implementation Fault is 86 of 125 (68.8%), where the agent over-trusts a topically matched skill and treats its reusable defaults, examples and templates as task-specific requirements. |  |
| B137 | p.8 | Agent Skills Can Be Harmful found 62.6% of efficiency regressions are Excessive Procedure, skills converting optional verification checklists into mandatory work, with Excessive Verification alone at 36.8%, rather than prompt length. |  |
| B142 | p.7-8 | Agent Skills Can Be Harmful traced context-overhead cases to 43 of 46 from the always-loaded mandatory skill body versus 3 from lazy-loaded supplementary material. |  |
| B194 | p.10 | Agent Skills Can Be Harmful draws 307 confirmed failures from a deliberately expanded 20,664-pair space, so the 307 figure carries no prevalence meaning, and the paper's external-validity note goes no further than saying its findings may not directly apply to all settings. | "may not directly apply to all settings" |
| C032 | p.5-6 | Applicability Mismatch, the wrong-skill case everyone worries about, is 2 of 125 confirmed functional failures (1.6%). |  |
| C033 | p.5-6 | Task-Implementation Fault is 68.8% of confirmed functional failures, with the mechanism being over-trust of a topically matched skill's reusable defaults, examples and templates as task-specific requirements. |  |
| C056 | p.8 | 62.6% of efficiency regressions are Excessive Procedure, skills turning optional verification into mandatory work, against 25.3% context bloat. |  |
| C062 | p.7-8 | 43 of 46 context-overhead cases came from the always-loaded body and 3 from supplementary material. |  |
| C091 | p.7-8 | 43 of 46 context-overhead cases came from the mandatory body. |  |

## Towards a Risk Assessment of Malicious Skill Files in Coding Agents

arXiv [2608.05223](https://arxiv.org/abs/2608.05223) · 10 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B160 | p.16 | Given delegated auto-approved privileges, Gemini CLI was exploited in 96.1% of 2,816 runs and Qwen Code in 74.0% of 2,813 across 2,826 synthetic malicious skill files. |  |
| B161 | p.20 | Explicit refusal appeared in 1.99% of all 5,629 runs. |  |
| B162 | p.22, p.24 | For Gemini, non-exploited runs were mostly inattention: 45.0% acknowledged-but-didn't-execute and 39.6% no-acknowledgment, against 15.3% recognised-attack. |  |
| B163 | p.22, p.24 | Even the least exploitable ATT&CK tactic, exfiltration, pooled at 67.2%. |  |
| B164 | p.22, p.24 | Exploitability was flat across six different generator models, including small open-weight ones, so a low-resource local attacker suffices. |  |
| C068 | p.16, p.20 | Agents with delegated privileges complied with malicious skill files in 95.5-96.1% of runs for Gemini CLI and 71.6-74.0% for Qwen Code depending on the estimator. |  |
| C069 | p.16, p.20 | Explicit refusal occurred in 1.99% of all 5,629 runs. |  |
| C076 | p.25 | The Transurban authors' own recommendation is to treat third-party skills as unverified binaries. |  |
| C080 | p.3, p.16, p.18 | Three judges scoring the same runs reached Fleiss' kappa = -0.06, i.e. chance, on one agent against 0.51 on the other. |  |
| C081 | p.3, p.16, p.18 | Individual flag rates on the chance-agreement agent were 94.8%, 72.3% and 22.4%, so a lone judge could misstate the result by more than 20 points. |  |

## From Context to Skills: Can Language Models Learn from Context Skillfully?

arXiv [2604.27660](https://arxiv.org/abs/2604.27660) · 9 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B003 | p.7 | Ctx2Skill raised GPT-4.1 from 11.1% to 16.5% and GPT-5.1 from 21.1% to 25.8% over 1,899 tasks with no error bars. |  |
| B064 | Eq. 2, p.4 | In Ctx2Skill the raw context C is always supplied alongside the generated SKILL.md at inference. |  |
| B065 | p.7 | In Ctx2Skill, even with both context and skill, the best configuration solves 25.8% of CL-bench, leaving about 74% unsolved. |  |
| B112 | p.8, p.19 | Ctx2Skill finds GPT-4.1's fixed-iteration performance falls non-increasingly across five self-play iterations, 15.9 -> 15.6 -> 15.6 -> 15.2 -> 14.7%, while median skill length grows from 311 to 1,703 words. |  |
| B113 | p.8, p.19 | In Ctx2Skill, selecting short early skill sets recovers GPT-4.1 performance to 16.5%. |  |
| B114 | p.19 | In Ctx2Skill, GPT-5.1's median skill length growth runs from 1,235 to 6,447 words. |  |
| B115 | p.16 | In Ctx2Skill, GPT-5.2's own self-play solved rate falls from 36.1% to 23.0%. |  |
| B116 | p.6 | Ctx2Skill's authors call this adversarial collapse and note it is undetectable inside the loop because each iteration's judge only sees that iteration's new tasks. | "adversarial collapse" |
| C016 | Eq. 2, p.4 | Ctx2Skill always supplies raw context alongside the generated skill. |  |

## SkillCommit: Evolving Agent Skills through Behaviorally Validated Scope Expansion

arXiv [2608.15165](https://arxiv.org/abs/2608.15165) · 9 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B004 | p.5 | SkillCommit raised the mean task metric from 50.80% to 78.15% and was best on 18/18 pairs across 6 tasks x 3 configs. |  |
| B040 | p.5-6 | SkillCommit's help/hurt decomposition shows ACE repairs 3 and breaks 24 on RuleArena Airline, and Trace2Skill repairs 2 and breaks 24. |  |
| B041 | p.5-6 | SkillCommit finds 31 of 72 task-level results from four competing skill methods fall below the no-skill baseline. |  |
| B118 | p.4 | SkillCommit accepts no widening of a skill's claimed scope unless replay on every source instance still succeeds, using Commit / Split / Reject with previous versions preserved for traceability and all candidate versions plus replay outcomes retained for audit. |  |
| B119 | p.5-6 | SkillCommit repairs 24 and 21 failures on the two RuleArena subsets while regressing on 0 and 2 cases, against ACE's 3/24 and 3/13. |  |
| B120 |  | SkillCommit contains no cost accounting anywhere in the paper, and its skills are when/repair/avoid_when records on closed-ended benchmarks with automatic verifiers rather than human-authored SKILL.md files on open-ended work. |  |
| C035 | p.5-6 | SkillCommit finds 31 of 72 task-level results from four published skill methods fall below the no-skill baseline, with ACE repairing 3 and breaking 24 on one subset. |  |
| C037 | p.4-6 | SkillCommit's replay gate commits only if the widened skill still passes replay on every source instance, otherwise narrowing or rejecting, and it repairs 24 and 21 failures while regressing on 0 and 2. |  |
| C038 |  | SkillCommit reports no cost accounting at all. |  |

## Don't Offer What Can't Be Done: Deterministic Executability Gating for LLM Skill Selection at Scale

arXiv [2608.01050](https://arxiv.org/abs/2608.01050) · 9 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B050 | p.5 | In Wix Helpmate's live system, 59.4% of topically-relevant skill-message candidates were literally impossible to execute for that account. |  |
| B055 | p.5 | Wix inverted each skill's own hard exit conditions into a deterministic precondition gate, dropping 1,039,462 of 1,749,270 candidate pairs (59.4%) and 59.1% of skill-description tokens, cutting skill context by 90.5% overall against exposing all ten skills. |  |
| B056 | p.5 | With the gate removed, Wix's model selected a would-be-blocked skill in 78 of 1,000 replayed conversations, in a deliberately risk-enriched cohort the authors say must not be extrapolated to all chatbot messages. | "must not be extrapolated to all chatbot messages" |
| B057 | p.6 | Wix's result covers one ten-skill topic family, and at Scale in the title means 756K messages, not many skills. |  |
| B141 | p.4-5 | Wix reaches a 90.5% reduction the same way, with semantic matching removing 76.9% of the all-message baseline first and the executability gate then removing 59.1% of what remains. |  |
| C029 | p.4-5 | Wix found 59.4% of topically-relevant candidates were literally not executable for that account. |  |
| C030 | p.4-5 | Wix's semantic matching plus a deterministic precondition gate cut skill-description context by 90.5% combined, the gate's own share being 59.1% of post-semantic tokens, and stopped the model selecting a blocked skill in 78 of 1,000 risk-enriched replayed conversations. |  |
| C060 | p.4-5 | Semantic matching plus precondition gating gives a 90.5% combined reduction. |  |
| C092 | p.4-5 | 59.4% of topically-relevant candidates were unrunnable, and semantic matching plus gating cut skill-description context 90.5% combined, with the gate contributing 59.1% of that. |  |

## @skills: Attention is all you have

arXiv [2608.12610](https://arxiv.org/abs/2608.12610) · 8 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B049 | p.3, p.16 | @skills argues, without measuring it, that 56,804 indexed skills compete for fewer than 100 reliable auto-trigger slots. |  |
| B066 | p.2, p.13, p.26 | @skills scopes a skill to procedural knowledge, stating that a model knows what the world knows and a skill adds what it doesn't. | "A model knows what the world knows; a skill adds what it doesn't" |
| B067 | p.2, p.13, p.26 | @skills's related work states that RAG retrieval is implicit and similarity-based while @skills references are explicit and deterministic, which is what reliability requires for instructions as opposed to facts. | "RAG retrieval is implicit and similarity-based; @skills references are explicit and deterministic, which is what reliability requires for instructions as opposed to facts" |
| B068 | p.24-25 | Nothing in the @skills protocol addresses episodic memory, user state or history, and discovery is named as its open gap. |  |
| B131 | p.14 | @skills argues for reference-based delivery on attention grounds, runs no controlled experiment, and its own recommendation is a small saved set rather than a large on-demand catalogue. |  |
| B180 |  | @skills makes an architectural argument that progressive disclosure is the answer. |  |
| C017 | p.26 | The @skills authors state that skills are explicit and deterministic, which is what reliability requires for instructions as opposed to facts. | "explicit and deterministic, which is what reliability requires for instructions as opposed to facts" |
| C028 | p.3, p.16 | @skills argues rather than measures that the realistic ceiling on an auto-triggered skill library is order a few hundred, plausibly under a hundred for reliable auto-trigger. |  |

## What Keeps Agent Skills from Being Reusable? Evidence from 138K SKILL.md Files

arXiv [2608.08453](https://arxiv.org/abs/2608.08453) · 8 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B051 | p.4 | In a 138K SKILL.md corpus, 52.3% of public skills have no trigger guidance and 13.5% have non-functional descriptions. |  |
| B146 | p.4 | Of 138,133 deduplicated public skills, 89.3% trigger at least one official-spec detector and 91.8% at least one defect of any tier, mean 2.5 per skill, with the headline threshold-robust at 88.8-94.6%. |  |
| B147 | p.4 | Top defects are missing trigger guidance at 52.3%, name-as-heading duplication at 44.3% and too many inline examples at 32.1%. |  |
| B148 | p.6 | Defect density scales with size (Spearman rho = 0.508): skills over 500 lines average 4.76 defects against 1.48 for skills under 50 lines. |  |
| B149 | p.4, p.6 | Self-marked AI-authored skills average 3.23 versus 2.34 defects, with safety defects 2.3x and portability defects 2.8x higher, which the authors say is associational rather than causal. |  |
| C064 | p.4 | 89.3% of public skills violate the official spec. |  |
| C094 | p.5 | Spec-aware descriptions correlate with 1.83 defects versus 3.00, Cliff's delta = -0.40, which is partly definitional since the detectors derive from the same spec. |  |
| C095 | p.4 | 52.3% of public skills have no trigger guidance at all. |  |

## When Experience Becomes Instruction: Trajectory Poisoning in Self-Evolving Agent Skill Systems

arXiv [2608.05563](https://arxiv.org/abs/2608.05563) · 8 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B165 | p.1, p.5-6 | In trajectory poisoning, a contributor supplying 3 records inside a 30-record batch (10% support) got an attacker-chosen behaviour written into the persistent skill in 546/600 trials (91.0% SER) across six evolvers, transferring to a structurally different system at 61.5%. |  |
| B166 | p.1, p.5-6 | Recurrence is the lever rather than volume: 5/25 at k=1, 21/25 at k=2, 25/25 at k=3, and still 22/25 when diluted to 3% support. |  |
| B167 | p.6 | The poisoned skill still outperforms no-skill on benign tasks. |  |
| B168 | p.7 | The trajectory-poisoning paper states that authentic provenance is not trustworthy provenance, because the artifact genuinely came from the victim's own evolver. | "Authentic provenance is not trustworthy provenance" |
| B169 | p.7 | A provenance-diversity gate requiring independent multi-cluster support cut post-gate success from 25/25 to 0/25. |  |
| C071 | p.5 | Three ordinary-looking records inside a batch of thirty got an attacker-chosen behaviour written into the persistent skill in 91.0% of 600 trials. |  |
| C072 | p.7 | The resulting poisoned skill was still no worse than no-skill on a 100-task benign split (20.0% versus 18.0% Hard, no significance test), and its provenance was genuine because the victim's own evolver wrote it. | "Authentic provenance is not trustworthy provenance" |
| C075 | p.7 | A pilot provenance-diversity gate requiring independent multi-cluster support before evidence is promoted took 25/25 to 0/25 on a single behaviour family at n = 25 with one benign control, which the authors call preliminary while noting coordinated Sybils could mimic diversity. |  |

## Skill Use or Skill Theater? Evaluating the Reasoning Backroom in Skill-Augmented Language Agents

arXiv [2607.27484](https://arxiv.org/abs/2607.27484) · 7 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B013 | p.6 | Skill Use or Skill Theater? finds delta-Acc within +/-.05 for the large majority of cells. |  |
| B014 | p.6 | Skill Use or Skill Theater? finds a misleading skill costs GPT-5-mini -.19. |  |
| B042 | p.13, p.6 | Skill Theater shows that under content-swapped skills Qwen2.5-7B makes 42 helpful and 46 harmful flips and DeepSeek-R1-7B 56 and 67, with near-zero delta-Acc concealing large decision churn. |  |
| C036 | p.13 | Skill Theater shows flat averages concealing 42 helpful and 46 harmful flips in the same condition. |  |
| C077 | p.7 | Four visible-signal detectors on 3,600 runs scored precision .31-.37 against a .34 base rate, at or below chance. |  |
| C078 | p.5 | Stated skill use sits at .88-1.00 across all twelve models while causal reliance ranges .00-.68. |  |
| C079 | p.6, p.14 | In multi-agent teams with no skill supplied at all, false provenance attribution ran at 1.00 for eight of twelve systems. |  |

## SkillsMetric: Mapping the Detection Boundary of Static Analysis for Malicious Agent Skills

arXiv [2608.08468](https://arxiv.org/abs/2608.08468) · 7 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B153 | p.2-3 | SkillsMetric reports recall that never exceeds about 62% under either protocol, with host destruction 0.0% (N=11), environment manipulation 22.2%, remote script 27.6%, split module 33.3% and prompt injection 41.7%. |  |
| B154 | p.2-3 | SkillsMetric reports supply chain detection at 100%, data exfiltration 92.9% and steganography 92.9%, so structural artefacts get caught while semantically normal code used for abnormal purposes does not. |  |
| B155 | p.3, p.5 | SkillsMetric's population scan flags 1.75% of 138K skills, explicitly a lower bound because only SKILL.md text was scanned and malicious logic typically resides in companion scripts. | "malicious logic typically resides in companion scripts" |
| B156 | p.3, p.5 | SkillsMetric's flagged tail on inspection was mostly dual-use tooling, not malice. |  |
| C065 | p.2-3 | One representative static framework reaches 62.2% recall at its F1-optimal operating point, with 89.4% precision and AUC 0.93. |  |
| C066 | p.2-3 | That static framework scores 0% on host destruction and 41.7% on prompt injection. |  |
| C070 | p.1, p.5 | SkillsMetric states that skill runners typically inject SKILL.md content as user-level messages, granting third-party content the same authority as direct user commands, a structural vulnerability that no amount of static analysis can address. | "skill runners typically inject SKILL.md content as user-level messages, granting third-party content the same authority as direct user commands... even a well-intentioned agent cannot distinguish skill instructions from user intent - a structural vulnerability that no amount of static analysis can address" |

## SkillSight: Calibrating Generic Content Bias for Skill Retrieval

arXiv [2607.18785](https://arxiv.org/abs/2607.18785) · 6 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B008 | p.6 | SkillSight reports end-to-end gains of +1.84 to +4.97 points over LLM selection across 3 agent models. |  |
| B048 | p.3 | SkillSight reports mean share of corpus ranked above the gold skill of 1.24% on SRA-Bench versus 0.42% SciFact, 0.23% FiQA and 0.05% ArguAna, which at 26k candidates is about 320 documents ahead of the right one. |  |
| B129 | p.6 | SkillSight found progressive disclosure the worst baseline end to end, with Overall 31.91 / 65.91 / 53.79 across three models, below plain LLM selection. |  |
| B130 | p.6 | In SkillSight, progressive disclosure collapses to 0.35 on a 0-100 scale on one task, in a configuration averaging 74.18s latency and 10,942 input tokens across all six tasks. |  |
| B182 |  | SkillSight finds progressive disclosure the worst baseline, with one task collapsing to 0.35. |  |
| C027 | p.3 | SkillSight finds 1.24% of the corpus ranks above the gold skill on average versus 0.05-0.42% for standard IR benchmarks, because skill documents share heavy boilerplate. |  |

## When and How Context Rot Appears in Coding Agents: A White-Box Study of Agent Skills in Code Auditing

arXiv [2607.17937](https://arxiv.org/abs/2607.17937) · 6 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B075 | p.7 | Context Rot holds a fixed 24-check audit skill constant and finds a 10,991-character clean context passes 8/10 runs while two different 299,140-character contexts pass 3/10 each. |  |
| B076 | p.8 | In Context Rot, coverage retention stays high at 0.93-0.95 while strict-success retention falls to 0.375. |  |
| B077 | p.1 | Context Rot states that rot often removes a few decisive obligations rather than the whole artifact. | "rot often removes a few decisive obligations rather than the whole artifact" |
| B078 | p.9 | In Context Rot, 38 of 44 failed runs (86.4%) end with an explicit success or completion claim, in a mixed inventory the authors say is explicitly not a prevalence estimate. | "not a prevalence estimate" |
| B079 | p.9 | Context Rot's working fix was re-stating the 24 obligations at the completion boundary, giving 10/10 versus 5/10 for a generic self-check, Fisher p = 0.0325. |  |
| C097 | p.9 | Re-stating critical obligations at the completion boundary gives 10/10 versus 5/10 for a generic self-check, p = 0.0325. |  |

## SkillEval: Decomposing Agent Skill Quality into Interpretable Signals

arXiv [2608.06891](https://arxiv.org/abs/2608.06891) · 6 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B151 | p.5-6 | SkillEval's learned document-level quality directions correlate with observed uplift at Pearson r = 0.779-0.787. |  |
| B152 | p.5-6 | SkillEval's metric-guided revision lifted mean downstream pass rate from 18.6% to 48.1% versus 38.3% for blind LLM revision. |  |
| B176 |  | SkillEval reports +29.5 pp from guided revision, supporting the position that curating procedure helps. |  |
| B178 |  | SkillEval reports quality scores predicting utility at r approximately 0.78. |  |
| C086 | p.5-6 | SkillEval learns quality directions supervised against downstream uplift and reaches r approximately 0.78. |  |
| C087 | p.5-6 | SkillEval's metric-guided revision lifts pass rate from 18.6% to 48.1% against 38.3% for blind revision. |  |

## ColluSkill: Adversarial Cross-Skill Composition for Evading Agent Skill Scanners

arXiv [2608.09732](https://arxiv.org/abs/2608.09732) · 6 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B157 | p.5 | ColluSkill splits a payload across three individually plausible skills and reports 96.0% average attack success across six deployed scanners (CISCO 100%, SkillFortify 100%, SkillSpector 99.0%, SlowMist 93.5%, Vetter 92.0%, Auditor 91.5%) against 14.2-39.0% for single-skill baselines. |  |
| B158 | p.7 | In ColluSkill, splitting alone gets only 36.7%, chain planning lifts it to 68.2%, and scanner-feedback rewriting takes it to 96.0%, so any scanner returning actionable feedback is a free oracle. |  |
| B159 |  | ColluSkill's context-aware defence ChainGuard cuts attack success to 22.5% at 99.5% benign pass, and the authors concede it does not reduce the ASR to zero. | "does not reduce the ASR to zero" |
| B193 |  | ColluSkill shows composition is exactly where the security model breaks. |  |
| C067 | p.2, p.5, p.7 | A payload split across three plausible skills, chain-planned and then iteratively refined against scanner feedback, reaches 96.0% average attack success across six scanners, while naive splitting alone gets only 36.7%, so the mechanism is adaptive refinement, not decomposition. |  |
| C074 | p.5, p.7 | ColluSkill's context-aware scanner, reading a candidate against everything already installed, takes attack success from 96.0% to 22.5% ASR at 99.5% benign pass, though it was never tested against an attacker adapting to it. |  |

## Comparative Approaches to Agent Retrieval over Large Skill Libraries

arXiv [2608.06196](https://arxiv.org/abs/2608.06196) · 5 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B047 | p.5 | Praetorian's production library of 690 skills achieves hybrid hit@5 = 0.735 and top-1 = 0.504, with indirect queries at 0.628 and about 26.5% of realistic queries unserved. |  |
| B140 | p.4 | Praetorian reports a whole catalogue at 46,915 tokens/task versus on-demand retrieval at about 560 tokens, a 98.8% reduction, with the authors saying the saving comes from loading less rather than from loading smarter. | "from loading less rather than from loading smarter" |
| C026 | p.5 | A production library of roughly 640-690 skills (875-entry catalogue) retrieves the right skill in the top 5 only 73.5% of the time, 62.8% on indirect queries, leaving roughly a quarter of realistic queries unserved. |  |
| C059 | p.4 | Whole catalogue loading costs 46,915 tokens versus on-demand at about 560. |  |
| C085 | p.5 | Retrieval evals are inflated by about 20 points when authors write the queries. |  |

## GitSkills: A Dataset of Agent Skills on GitHub

arXiv [2608.10906](https://arxiv.org/abs/2608.10906) · 4 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B143 | p.1-2 | GitSkills mined 3,797,117 SKILL.md files from 282,200 repositories and 195,841 accounts in July 2026, nine months after the format appeared, collapsing to 1,877,981 distinct contents, i.e. 50.5% verbatim copies. |  |
| B145 | p.1 | GitSkills states that skills have no central registry or package manager, so they spread by copying folders between repositories. | "have no central registry or package manager, so they spread by copying folders between repositories" |
| B188 | p.2 | GitSkills lists as an open question how often skills become outdated relative to the projects and tools they describe. | "How often do skills become outdated relative to the projects and tools they describe?" |
| C063 | p.1-2 | There are 3.8 million SKILL.md files, no registry, and 50.5% verbatim copies. |  |

## Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents

arXiv [2608.12851](https://arxiv.org/abs/2608.12851) · 4 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| B170 | p.5 | Practice Makes Unsafe finds that across 21 evolved configurations on a fixed backbone, all 21 author unsafe artifacts, 19 retrieve them, 19 show benign-task contamination, and 15 still cause harm in a clean session that reloads only the exported SKILL.md. |  |
| B171 | p.5 | In Practice Makes Unsafe, benign utility beats the no-evolution control in 15 of 21 settings while malicious success rises in 17 of 21. |  |
| B172 | p.6-7 | In Practice Makes Unsafe, three malicious tasks are enough: pooled carryover attack success goes 16.0% -> 35.3% after one three-task exposure and 41.3% at full budget, and benign-heavy update batches do not wash it out (31.8% versus 34.2% contamination for fully-mixed versus batched schedules). |  |
| C073 | p.5 | Practice Makes Unsafe reports all 21 evolved configurations authored unsafe artifacts and 15 still caused harm in a clean session that loaded only the exported SKILL.md. |  |

## Authoring Agent Skills: A Software-Engineering Approach

arXiv [2607.25032](https://arxiv.org/abs/2607.25032) · 1 claims

| ID | Page | Claim | Quotation |
|---|---|---|---|
| C096 | p.8 | A standing instruction in a skill body, however firmly worded, is something the model may read, defer, or skip, and only a hook can block. | "A standing instruction in a skill body, however firmly worded, is something the model may read, defer, or skip" |

## Cross-corpus claims

Claims about the corpus as a whole, or about the absence of evidence, which no single paper supports.

| ID | Claim |
|---|---|
| B184 | Ecosystem papers report 50-60% duplication, arguing dedup is essential. |
| C031 | No paper runs a controlled 10 -> 100 -> 1,000 -> 10,000 library-size sweep. |
