Daily brief · 6 stories
Wednesday, 7 October 2026
Most of today's useful papers are about whether our measurements can be trusted, which is a good lens for LLM security and time series alike.
Oracles and attack surfaces. The reproducibility audit of LLM/agent vulnerability-validation work is the one to read. Of 104 papers, only 59 had reachable artifacts. In 58 of 102 anchor cases the internal CVE ID disagreed with the directory label. The embedded oracles were weak: 20 of 30 patched counterfactuals still fired, and 7 of 19 negative controls triggered, giving 60% sensitivity and 45% specificity. A vulnerable-build trigger without a clean patched counterfactual isn't CVE-specific evidence. That makes me want a harder look at CyberFactory, which builds executable, verifiable tasks from public CVEs and reports a 22.8-point Pass@1 gain on CyberGym. Its verification quality is the thing to check.
On the agent-skills front, ElasticBack is a weight-free backdoor that needs both a malicious rule in the skill document and a benign-looking trigger in the user query. It reports near-zero false positives and evades deployment-time defenses. SkillsMetric maps the defender's side: static analysis hits AUC 0.93 and 93% on exfiltration, but catches 0% of host destruction via common shell commands and 42% of prompt injection. Together they say skill supply chains need semantic review, not just pattern matching. The DeepSeek Harness indirect-injection study (14,560 runs, 16 channels) is consistent with this. Hidden Unicode in files reached 25.5% ASR and the skills channel 16.0% under the rule-based judge. The two judges also disagreed on partial compliance, 7.3% versus 2.0%, so judge choice matters as much as the attack.
Internal signals. Two papers pull in opposite directions. The latent-safety-probe reproduction matches the original within 0.37 F1 on LLaMA-3.1-8B, and simple MLP probes land within about one point on Gemma, M
Today's stories
-
From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts
This pre-registered reproducibility audit examines LLM/agent-driven vulnerability validation artifacts across 104 papers (2023–2026), finding only 59 with publicly reachable artifacts and executing 18 paper-level workflows and 102 anchor benchmark cases. It reports high mismatch and fragility: 58/102 anchor cases have internal CVE IDs diverging from directory labels; only 10/18 artifacts run end-to-end at R0 (11/18 after R1 env-only fixes); and embedded oracles are unreliable, with 20/30 patched-counterfactuals still signaling and 7/19 negative controls triggering, yielding oracle sensitivity 60% and specificity 45%. The study argues that vulnerable-build triggers are not CVE-specific evidence without a clean patched counterfactual and offers a reusable, pre-registered protocol (post-conditions, R0/R1 repair ladder, G1–G3 evidence levels, patched-counterfactual oracles) for the community.
Why it matters It surfaces concrete failure modes and a reusable protocol for verifying LLM/agent security claims—directly informing eval design, oracle construction, and counterfactual controls for both LLM security and time-series-style benchmark reproducibility.
-
ElasticBack: Stealthy Conditional Backdoor in LLM-Agent Skills via Coupled Trigger-Rule Optimization
The paper introduces ElasticBack, a conditional single-skill backdoor for LLM agents that couples a malicious rule R in the skill document with a benign-looking trigger T in user queries, activating only when both appear. It constructs R via semantic-anchored rule injection and then fixes R while evolving T using a stealth-constrained genetic search to optimize attack effectiveness and stealth without modifying model weights. Experiments over three target behaviors (50 skills each) and four agent LLMs show high ASR with near-zero FPR, preserved clean accuracy, cross-model transfer, and evasion of deployment-time defenses, motivating stronger supply-chain defenses for skills.
Why it matters Highlights a practical, weight-free, conditional backdoor vector in the emerging agent skill ecosystem, relevant to securing tool-using LLMs and evaluating defenses against stealthy trigger-rule couplings.
-
Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection
This work evaluates indirect prompt injection against DeepSeek Harness using AI-Infra-Guard to generate tests, deliver controlled taint across 16 channels and two carrier modes, run 14,560 executions, and analyze traces with both rule-based and LLM-based judges. The strongest observed attack success rates are 17.0% under LLMJudge for a fake-completion attack (text), 25.5% under RuleJudge for hidden Unicode (file), and 16.0% under RuleJudge for the skills channel (file), with LLMJudge assigning more partial compliance than RuleJudge (7.3% vs. 2.0%). The study connects outcomes to DSH’s handling of tool results, extra contexts, and tool-call policy hooks, and proposes controls between untrusted content and sensitive actions, with code released at the linked repository.
Why it matters Provides a large-scale, reproducible, channel- and payload-stratified benchmark of indirect prompt injection on an agentic LLM stack with dual judges, informing defense design and evaluation baselines.
-
LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models
The paper introduces LiveHouse-TS, an open-world, living benchmark infrastructure for evaluating TSFMs via prequential testing on real future data rather than static windows. It shifts evaluation focus from snapshot accuracy to continuous temporal validity in environments with seasonality, distribution shifts, and unexpected events, aiming to answer long-term questions about ranking stability and robustness. Streaming evaluations across 11 domains and 17 datasets show that model rankings from static benchmarks dramatically reshuffle under the live protocol.
Why it matters If you work on LLM security or TSFMs, this exposes how static benchmarks can mislead robustness assessments under shifts, offering an infrastructure to stress-test models’ long-term validity—analogous to live red-teaming and continual evals in LLMs.