By section · Wednesday, 7 October 2026
Everything that cleared the bar
The full catalog behind today's front page — 11 of 15 sections have something in them. A story can top its section and still miss the flat front-page cut, so this is not the leftovers.
Research
LLM security
Attacks, defenses, and evaluation of LLM systems: jailbreaks, prompt injection, data extraction, red-teaming, alignment-adjacent safety failures with a security framing.
-
From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts
This pre-registered reproducibility audit examines LLM/agent-driven vulnerability validation artifacts across 104 papers (2023–2026), finding only 59 with publicly reachable artifacts and executing 18 paper-level workflows and 102 anchor benchmark cases. It reports high mismatch and fragility: 58/102 anchor cases have internal CVE IDs diverging from directory labels; only 10/18 artifacts run end-to-end at R0 (11/18 after R1 env-only fixes); and embedded oracles are unreliable, with 20/30 patched-counterfactuals still signaling and 7/19 negative controls triggering, yielding oracle sensitivity 60% and specificity 45%. The study argues that vulnerable-build triggers are not CVE-specific evidence without a clean patched counterfactual and offers a reusable, pre-registered protocol (post-conditions, R0/R1 repair ladder, G1–G3 evidence levels, patched-counterfactual oracles) for the community.
Why it matters It surfaces concrete failure modes and a reusable protocol for verifying LLM/agent security claims—directly informing eval design, oracle construction, and counterfactual controls for both LLM security and time-series-style benchmark reproducibility.
-
ElasticBack: Stealthy Conditional Backdoor in LLM-Agent Skills via Coupled Trigger-Rule Optimization
The paper introduces ElasticBack, a conditional single-skill backdoor for LLM agents that couples a malicious rule R in the skill document with a benign-looking trigger T in user queries, activating only when both appear. It constructs R via semantic-anchored rule injection and then fixes R while evolving T using a stealth-constrained genetic search to optimize attack effectiveness and stealth without modifying model weights. Experiments over three target behaviors (50 skills each) and four agent LLMs show high ASR with near-zero FPR, preserved clean accuracy, cross-model transfer, and evasion of deployment-time defenses, motivating stronger supply-chain defenses for skills.
Why it matters Highlights a practical, weight-free, conditional backdoor vector in the emerging agent skill ecosystem, relevant to securing tool-using LLMs and evaluating defenses against stealthy trigger-rule couplings.
-
Learning the Pareto Frontier of Predictive Models under Distribution Shift
The paper introduces Frontier Learning, which treats a library of pretrained models with mixed access regimes (black-box predictions and white-box representations) as complementary sources under distribution shift. It builds a unified target-domain feature by concatenating internal representations from white-box models with outputs from black-box models, then trains a lightweight regularized supervised learner, guaranteeing training-sample risk no worse than any single baseline including zero-shot reuse, fine-tuning, or direct training. Experiments on simulations and real-world shifts (DomainNet/VisDA and MIMIC-IV-Notes ICU mortality) show it matches or outperforms the strongest individual reuse strategy, with the largest gains when no single baseline is reliable across the shifts considered.
Why it matters Provides a simple, access-agnostic ensemble-on-representations approach that subsumes common reuse strategies and is empirically strong under shift—relevant to robustness/security and modular reuse of foundation models.
-
Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection
This work evaluates indirect prompt injection against DeepSeek Harness using AI-Infra-Guard to generate tests, deliver controlled taint across 16 channels and two carrier modes, run 14,560 executions, and analyze traces with both rule-based and LLM-based judges. The strongest observed attack success rates are 17.0% under LLMJudge for a fake-completion attack (text), 25.5% under RuleJudge for hidden Unicode (file), and 16.0% under RuleJudge for the skills channel (file), with LLMJudge assigning more partial compliance than RuleJudge (7.3% vs. 2.0%). The study connects outcomes to DSH’s handling of tool results, extra contexts, and tool-call policy hooks, and proposes controls between untrusted content and sensitive actions, with code released at the linked repository.
Why it matters Provides a large-scale, reproducible, channel- and payload-stratified benchmark of indirect prompt injection on an agentic LLM stack with dual judges, informing defense design and evaluation baselines.
-
When Do PEFT Adaptations Leak Structure? Measuring Black-Box Structural Bounds in Public-Base Model Services
The paper introduces VectorHijack-SR, a black-box measurement method that turns paired victim/base residuals into calibrated bounds over PEFT family, layer locality, and coarse rank, using aggregated query-level statistics and a service-disjoint classifier plus a cross-fitted hierarchical rejector to test LoRA-manifold membership. Empirically, family leakage exceeds chance across multiple backbones/tasks, rank inference varies by task, the rejector attains AUROC 0.804 and high known-set accuracy but struggles with structurally close DoRA/LoRA+head, and exact-version linkage on held-out LoRA-r64 services reaches AUC 0.940. Despite measurable structural leakage, experiments show a visibility–exploitability gap: two-stage recovery offers no fair-budget query savings, posterior-selected PEFT underperforms distill-then-convert PEFT, and free-running generation remains near chance.
Why it matters It provides concrete, black-box evidence and metrics for structural leakage from PEFTed services with known bases, informing red-team audits, model fingerprinting, and security evaluations of adapter deployment choices.
-
Deployable Per-Instance Multi-Layer Activation Steering for Large Language Models
The paper shows that in activation steering, the optimal injection layers vary per input, with per-instance multi-layer selection outperforming any fixed global layer set across six binary persona traits on two 8B models. A greedy layer-ranking rule based on single-layer marginal effects nearly matches a per-instance oracle but requires gold labels, so the authors train a prompt-only predictor to mimic it and propose a deployable recipe: a prompt-based per-instance layer ranker, a classifier to infer steering direction, and an adaptive gate that limits layers by short steered passes. This deployable approach recovers most of the oracle’s gains, avoids dropping below unsteered baselines on average, reduces fluency collapse relative to strong global selection, and is supported by a mechanistic account emphasizing direction over magnitude that explains misdirection flips, collapse from over-steering, and unsteerable ceilings.
Why it matters It operationalizes per-instance, label-free activation steering with adaptive multi-layer control, offering a practical path to safer, targeted behavior edits without sacrificing fluency—relevant to red-teaming defenses and controllable generation, and conceptually aligned with direction-based causal hypotheses familiar in interpretability.
-
SkillsMetric: Mapping the Detection Boundary of Static Analysis for Malicious Agent Skills
This work introduces SkillsMetric, a five-stage static analysis framework that scores LLM agent skills across pattern density, statistical anomaly, dataflow taint, import anomaly, and capability mismatch. Using an adversarial dataset of 2,266 skills covering 16 attack types and the SkillMD-138K corpus, it reports AUC 0.93 and 5-fold CV F1 73.4%±0.5%, with 93% detection for data exfiltration and steganographic payloads. It also exposes blind spots—0% detection for host destruction via common shell commands and 42% for prompt injection—arguing static analysis alone is insufficient and should be paired with semantic review in defense-in-depth.
Why it matters Directly maps the detection boundary of static analysis for agent skills, highlighting failure modes (e.g., prompt injection, host destruction) that inform red-teaming, hybrid detection pipelines, and security benchmarks for LLM agents.
-
Generating Attacks for LLMs with GFlowNets
The work proposes an automated, human-independent red teaming framework that trains an attacker LLM via GFlowNets to probe a specified victim LLM and yield a quantitative robustness score. It addresses limitations of manual expert red teaming and automated methods tied to fixed datasets by enabling adaptive, creative attack generation. The approach targets generating more effective English adversarial inputs than existing benchmarks and, uniquely, introduces Turkish-language attack generation.
Why it matters It explores GFlowNets for adaptive attack synthesis against LLMs, offering a path beyond fixed corpora and extending multilingual red teaming—a relevant direction for security evaluation and scalable adversarial data generation.
-
Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining
The paper introduces a task-agnostic measure of training data influence that quantifies how much an example’s gradient update reduces the squared distance to a run’s final parameters, estimated from intermediate checkpoints without retraining. Applied to 18 Pythia and PolyPythia configurations, the method reveals systematic temporal shifts in influential data: literature-related data align more with the parameter trajectory early, while STEM data align more later. This qualitative crossover is broadly consistent across configurations, offering a tractable trajectory-level perspective that complements task-based influence analyses.
Why it matters Gives a practical, retraining-free way to map which data sources steer pretraining trajectories over time, informing data curation, curriculum, and security-relevant provenance audits for LLMs.
-
Proving the Utility of Large Language Models in Cybersecurity Simulations: A Comprehensive Examination
The paper evaluates LLMs for cybersecurity simulations, using YAML to represent complex network configurations and to drive pipelines that support RL agent training. It compares LLM-based techniques to classical methods like Double Q-learning with PER, claiming improved efficiency, adaptability, and realism in cyberattack simulations. In benchmarks on multiple synthetic topologies, LLM-instantiated Python agents reached up to a 94.5% compromise rate and 0.02–0.06s per assessment, a ~25,000x–50,000x speedup over traditional RL cycles.
Why it matters Shows LLM-driven generation/execution pipelines can massively accelerate cyber-sim training and evaluation, relevant to secure LLM autonomy and sim-to-real robustness for time-series/agentic models.
-
FedLNS: Leverage LayerNorm Signature Modeling to Mitigate Adversarial Manipulation in Federated LLMs
The paper introduces FedLNS, a server-side framework that screens federated LLM client updates by modeling changes in trainable normalization-layer parameters as signatures and comparing them to a history-aware cross-client reference. It operates without additional client-to-server metadata, raw client data, trusted server datasets, labeled attack examples, or a separately trained detector, and retains full-model updates for standard or compatible aggregation after screening. Experiments on GPT-, BERT-, and LLaMA-style models trained from scratch with 200 clients show that under 40% target manipulation, FedLNS yields lower test perplexity than six baselines across IID and non-IID partitions.
Why it matters Offers a practical, data-free, model-internal signal for adversarial update filtering in federated LLM training, aligning with robustness and integrity concerns in LLM security.
-
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families
This study reproduces Khatri et al.’s pipeline for training lightweight MLP probes on final-layer activations to detect harmful prompts, matching the original LLaMA-3.1-8B results within 0.37 F1 points (0.2 on BeaverTails). It extends the evaluation across Gemma-4-E4B, Mistral-7B-v0.3, and Qwen2-7B on WildJailbreak, BeaverTails, and AEGIS 2.0, finding F1 within about one point of the LLaMA-3.1-8B values using the same probe architecture. It also assesses nondeterminism by varying seeds and reports that final token latent vectors were invariant to seed across architectures.
Why it matters Suggests simple latent-space probes generalize across model families with stable activations, informing scalable, low-overhead safety detection for LLMs and complementing guardrail strategies.
-
RouteGuard: Certifying Routing Gain in LLM Multi-Agent Systems When Complementarity Is Not Enough
The paper shows that common assumptions for multi-agent LLM routing—optimizing gate AUC and relying on advisor complementarity—do not determine deployable gain, and introduces RouteGuard to certify routing benefit. It decomposes gain as G = πΔ_E, identifies a conditional-regret functional Φ as the governing quantity (not AUC), provides a finite-sample certification bracket with a matching Le Cam lower bound that is constant-sharp over the fixed-activity class, and observes a robustness phase transition. Experiments on RouterBench and OpenRCA demonstrate RouteGuard acting as a guardrail, certifying gain only under prompt-level sampling (not workload-cluster resampling) on RouterBench and refusing to certify on OpenRCA due to advisor redundancy, with pre-registered semi-synthetic controls confirming calibration thresholds.
Why it matters It offers a principled, sample-sensitive certification of routing gain that can prevent illusory improvements in LLM multi-agent systems and aligns evaluation with deployment risk.
-
Measure, Don't Optimize: Forecasting Recovery in LLM Unlearning
The paper introduces J-Access, an inference-time audit that projects intermediate representations into vocabulary space to measure how often target concepts remain accessible along the output pathway, hypothesizing that residual accessibility predicts recovery susceptibility. Auditing 398 public unlearned models across eight methods, they find most retain accessibility above a retain-only gold level, pre-attack accessibility predicts recovery speed and extent at the model level (but not per-fact), and directly minimizing J-Access causes the model to hide knowledge from the audit while increasing post-attack recovery. They conclude J-Access is useful as a model-level diagnostic for residual susceptibility and caution against turning internal audits into optimization targets without validation.
Why it matters Offers an actionable, model-level risk signal for unlearning robustness and a clear warning about Goodharting internal audits—relevant for designing secure unlearning and audit pipelines in LLMs.
-
Stopping and Routing LLM Judge Panels
The work frames judge-panel design as a role-conditioned allocation problem that, using a small labeled audit set, declared slices, and judge costs, estimates target-relative roles: copies, complements, and specialists. These roles yield a policy to drop copies, add complements globally, route specialists on slices, and stop when validation gain falls below a threshold, producing a reusable, auditable call plan. It is empirically compared across multiple audit settings against single judges, flat panels, diversity heuristics, full-call stacking, reliability juries, and frugal cascades, yielding a regime map for when to route specialists, stop in saturated regimes, keep broad ensembles, and ignore conditional copies.
Why it matters It offers a concrete, auditable routing/early-stopping policy for multi-judge LLM evaluation that balances risk and cost across slices—directly relevant to secure deployment and panel design in LLM oversight and verifiers for time-series or reasoning tasks.
-
FutureBridge: Token Selection Beyond Local Preference in Collaborative Decoding
FutureBridge proposes a token reranking scheme for LLM–SLM collaborative decoding that prioritizes tokens enabling the SLM’s subsequent reasoning rather than the LLM’s local preference. Training uses an answer-verified LLM trajectory to fix a shared future, while a frozen SLM assigns counterfactual scores to candidate tokens under this common context to supervise a lightweight reranker observing only the current state and token. At inference, the LLM only expands the candidate pool and the reranker selects a single token before handing generation back to the SLM without appending any future suffix, yielding a 35.1% relative Math Avg. improvement for Qwen3-1.7B across five math reasoning benchmarks over greedy decoding.
Why it matters It reframes collaboration as SLM-aware token selection, offering a practical path to safer, stronger small-model reasoning with minimal LLM involvement—relevant to alignment, tool-use gating, and cost/security tradeoffs in cooperative decoding.
-
Diversity Matters: Distributional Feature Coverage Sample Selection for Data-Efficient Backdoor Attacks
The paper introduces DFCS, a training-free, trigger-agnostic sample selection method that clusters pretrained features into one region per poisoning slot and picks centroid-nearest samples to avoid redundant poisons. A local first-order analysis links this allocation to feature-coverage and representative-mass terms. On CIFAR-10, Tiny-ImageNet, and Imagenette with BadNets and Blended attacks, DFCS achieves the highest mean ASR among seven selectors in all six settings, averaging 96.30% and surpassing the strongest comparator by 4.60 points on average while maintaining clean accuracy.
Why it matters It highlights a simple, training-free selection principle that boosts low-budget dirty-label backdoor potency, informing both attack design and defenses that rely on data diversity and feature coverage.
-
ThreatLens: Evidence-Guided Ranking of High-Priority CVEs
ThreatLens is a deployment-realistic CVE prioritization framework that ranks at review time using only cutoff-valid evidence and learns exploitation relevance from future CISA KEV entries as weak supervision. Under forward-in-time, CVE-disjoint evaluation, it significantly outperforms CVSS, EPSS, and rule-based evidence fusion, surfacing 80.0% of future KEV CVEs in the top 20 and 95.9% in the top 50 on a held-out test split. Early-warning analysis shows it flags a substantial fraction of later KEV entries before catalog inclusion, enabling timely, evidence-grounded triage.
Why it matters Demonstrates a realistic, temporally valid ranking approach with weak supervision that materially beats common baselines—relevant to LLM security triage pipelines and time-aware evaluation of predictive models.
Adversarial ML
Adversarial examples, robustness, poisoning, and attacks/defenses on ML models generally (not LLM-specific).
-
ElasticBack: Stealthy Conditional Backdoor in LLM-Agent Skills via Coupled Trigger-Rule Optimization
The paper introduces ElasticBack, a conditional single-skill backdoor for LLM agents that couples a malicious rule R in the skill document with a benign-looking trigger T in user queries, activating only when both appear. It constructs R via semantic-anchored rule injection and then fixes R while evolving T using a stealth-constrained genetic search to optimize attack effectiveness and stealth without modifying model weights. Experiments over three target behaviors (50 skills each) and four agent LLMs show high ASR with near-zero FPR, preserved clean accuracy, cross-model transfer, and evasion of deployment-time defenses, motivating stronger supply-chain defenses for skills.
Why it matters Highlights a practical, weight-free, conditional backdoor vector in the emerging agent skill ecosystem, relevant to securing tool-using LLMs and evaluating defenses against stealthy trigger-rule couplings.
-
When Do PEFT Adaptations Leak Structure? Measuring Black-Box Structural Bounds in Public-Base Model Services
The paper introduces VectorHijack-SR, a black-box measurement method that turns paired victim/base residuals into calibrated bounds over PEFT family, layer locality, and coarse rank, using aggregated query-level statistics and a service-disjoint classifier plus a cross-fitted hierarchical rejector to test LoRA-manifold membership. Empirically, family leakage exceeds chance across multiple backbones/tasks, rank inference varies by task, the rejector attains AUROC 0.804 and high known-set accuracy but struggles with structurally close DoRA/LoRA+head, and exact-version linkage on held-out LoRA-r64 services reaches AUC 0.940. Despite measurable structural leakage, experiments show a visibility–exploitability gap: two-stage recovery offers no fair-budget query savings, posterior-selected PEFT underperforms distill-then-convert PEFT, and free-running generation remains near chance.
Why it matters It provides concrete, black-box evidence and metrics for structural leakage from PEFTed services with known bases, informing red-team audits, model fingerprinting, and security evaluations of adapter deployment choices.
-
SkillsMetric: Mapping the Detection Boundary of Static Analysis for Malicious Agent Skills
This work introduces SkillsMetric, a five-stage static analysis framework that scores LLM agent skills across pattern density, statistical anomaly, dataflow taint, import anomaly, and capability mismatch. Using an adversarial dataset of 2,266 skills covering 16 attack types and the SkillMD-138K corpus, it reports AUC 0.93 and 5-fold CV F1 73.4%±0.5%, with 93% detection for data exfiltration and steganographic payloads. It also exposes blind spots—0% detection for host destruction via common shell commands and 42% for prompt injection—arguing static analysis alone is insufficient and should be paired with semantic review in defense-in-depth.
Why it matters Directly maps the detection boundary of static analysis for agent skills, highlighting failure modes (e.g., prompt injection, host destruction) that inform red-teaming, hybrid detection pipelines, and security benchmarks for LLM agents.
-
Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining
The paper introduces a task-agnostic measure of training data influence that quantifies how much an example’s gradient update reduces the squared distance to a run’s final parameters, estimated from intermediate checkpoints without retraining. Applied to 18 Pythia and PolyPythia configurations, the method reveals systematic temporal shifts in influential data: literature-related data align more with the parameter trajectory early, while STEM data align more later. This qualitative crossover is broadly consistent across configurations, offering a tractable trajectory-level perspective that complements task-based influence analyses.
Why it matters Gives a practical, retraining-free way to map which data sources steer pretraining trajectories over time, informing data curation, curriculum, and security-relevant provenance audits for LLMs.
-
FedLNS: Leverage LayerNorm Signature Modeling to Mitigate Adversarial Manipulation in Federated LLMs
The paper introduces FedLNS, a server-side framework that screens federated LLM client updates by modeling changes in trainable normalization-layer parameters as signatures and comparing them to a history-aware cross-client reference. It operates without additional client-to-server metadata, raw client data, trusted server datasets, labeled attack examples, or a separately trained detector, and retains full-model updates for standard or compatible aggregation after screening. Experiments on GPT-, BERT-, and LLaMA-style models trained from scratch with 200 clients show that under 40% target manipulation, FedLNS yields lower test perplexity than six baselines across IID and non-IID partitions.
Why it matters Offers a practical, data-free, model-internal signal for adversarial update filtering in federated LLM training, aligning with robustness and integrity concerns in LLM security.
-
Diversity Matters: Distributional Feature Coverage Sample Selection for Data-Efficient Backdoor Attacks
The paper introduces DFCS, a training-free, trigger-agnostic sample selection method that clusters pretrained features into one region per poisoning slot and picks centroid-nearest samples to avoid redundant poisons. A local first-order analysis links this allocation to feature-coverage and representative-mass terms. On CIFAR-10, Tiny-ImageNet, and Imagenette with BadNets and Blended attacks, DFCS achieves the highest mean ASR among seven selectors in all six settings, averaging 96.30% and surpassing the strongest comparator by 4.60 points on average while maintaining clean accuracy.
Why it matters It highlights a simple, training-free selection principle that boosts low-budget dirty-label backdoor potency, informing both attack design and defenses that rely on data diversity and feature coverage.
-
ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions
ProbGuard reframes LLM safety assessment as probabilistic risk estimation over the model’s early output distribution rather than deterministic classification on completed token sequences. It estimates the probability of unsafe continuation via Monte Carlo over generated prefix distributions and calibrates this risk through post-training, enabling early stopping of unsafe generations. Empirically, it achieves the best calibration across nine model–dataset settings (average Brier and ECE reduced by 79.6% and 71.9% vs. the best baseline) and caps jailbreak attack success to ≤1% after only the first ten decoding steps.
Why it matters It operationalizes distribution-aware, early-stage safety control with strong calibration—useful for red-teaming defenses, risk-sensitive deployment, and aligning with probabilistic guardrail research and time-series-style forecasting over generation trajectories.
-
Analysis of Federated Aggregation under Model Poisoning and Backdoor Attacks: A Reconstructed Cross-Dataset and Cross-Architecture Benchmark
This work reconstructs a 500-cell federated aggregation benchmark across five methods, five datasets, five architectures, and four conditions (clean, sign-flipping, Gaussian, BadNets), recovering 454 original runs and 36 repairs/reruns, with 10 SVHN cells covered only by summary provenance. Trimmed Mean has the best clean macro-mean accuracy and lowest mean within-task rank, while Krum shows the top recorded accuracy under sign-flipping and Gaussian, and these rankings persist when restricted to fully logged task pairs. Audits reveal the BadNets metric is actually TTLR (all test inputs triggered) and a FedPARETO pathway where reported predictive summaries may not match the corrupted updates used for aggregation, so results are descriptive within recorded configs rather than universal robustness claims.
Why it matters Highlights reproducibility gaps, metric mis-specification (TTLR), and aggregation/provenance pitfalls that can confound robustness claims in federated settings relevant to secure FL and poisoning/backdoor evaluation.
-
Into the ORBIT for Time Series: Training Regimes for Foundation Models
The paper introduces ORBIT, a training paradigm for TSFMs that explicitly controls pre-training distributions via Bootstrap Multi-Level Sampling and Omni-Range Incremental Training. It trains Falcon-2.0, a univariate encoder-only Transformer with missingness-aware triple-channel patch tokenization and parallel patch prediction, under this regime. The authors also propose Rank-Guided Cross-Depth Alignment to align shallow and deep representations, and report strong zero-shot forecasting on GIFT-Eval and fev-bench across diverse domains and frequencies.
Why it matters It provides a principled, controllable pretraining regime and lightweight representation alignment that could transfer to secure LLM pretraining curricula and to robust TSFM deployment across heterogeneous, imbalanced time-series corpora.
-
Cross-Corpus Evaluation of Generalizable Vulnerability Detection in IoT Firmware
The paper introduces IoTVulBench, a human-verified, contamination-screened benchmark for cross-corpus IoT firmware vulnerability detection, and evaluates five model architectures, two tuning methods, and three curriculum strategies with ensemble, distillation, and robustness analyses. Models trained on IoTVulBench outperform matched single-source datasets (MCC 0.58 vs. 0.44 for PrimeVul and 0.39 for D2A), with staged curriculum learning raising MCC to 0.69 and a diversity-optimized ensemble reaching 0.73, surpassing a static analyzer (0.31) and PrimeVul by large margins while maintaining 86% performance under identifier renaming and showing strong calibration and largely faithful explanations. At a 0.5% FPR, the model missed 21% of vulnerabilities versus 71% for the strongest comparator, indicating domain-matched data and curriculum design—not model scale alone—drive generalization in firmware vulnerability detection.
Why it matters Provides a contamination-screened, human-verified benchmark and concrete training/curriculum/ensemble recipes that substantially improve cross-corpus generalization—useful for LLM security evaluation and for designing data- and curriculum-centric pipelines beyond scaling.
TSFM
Time series foundation models: pretraining, zero-shot forecasting, benchmarks, architectures for time series.
-
LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models
The paper introduces LiveHouse-TS, an open-world, living benchmark infrastructure for evaluating TSFMs via prequential testing on real future data rather than static windows. It shifts evaluation focus from snapshot accuracy to continuous temporal validity in environments with seasonality, distribution shifts, and unexpected events, aiming to answer long-term questions about ranking stability and robustness. Streaming evaluations across 11 domains and 17 datasets show that model rankings from static benchmarks dramatically reshuffle under the live protocol.
Why it matters If you work on LLM security or TSFMs, this exposes how static benchmarks can mislead robustness assessments under shifts, offering an infrastructure to stress-test models’ long-term validity—analogous to live red-teaming and continual evals in LLMs.
-
LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications
The paper surveys LLM-based forecasting agents, organizing architectures into standalone LLM workflows over temporal inputs, tool/retrieval-augmented agents, and hybrids with statistical or foundation models. It reviews training and evaluation practices and highlights negative evidence such as sensitivity to small perturbations, ablations showing no accuracy gains from the LLM, and potential benchmark contamination confounding temporal reasoning claims. Applications span finance, weather, health, energy, and operations, with the core limitation identified as measurement, motivating needs for robust calibration under shift, contamination-resistant live evaluation, explicit cost–accuracy reporting, and methods addressing forecast-induced feedback effects.
Why it matters Concise map of architectures, pitfalls, and eval gaps for LLM-in-the-loop forecasting—useful for security-minded eval design, contamination control, calibration under shift, and coupling with time-series FMs.
-
Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations
The authors test whether LLM hidden states during prefill encode vulnerability signals when reading C/C++ code by extracting last prefill token activations from four models and training small MLP probes. Across four function-level benchmarks (Devign, Big-Vul, Draper VDISC, PrimeVul), probes (13.4–16.0M params, <0.2% of base) reach 41.7% average F1, with a Qwen3.5-9B probe hitting 68.8% F1 on Devign, comparable to a fine-tuned-classifier SOTA (67.9%). Performance lags SOTA on the more imbalanced datasets, indicating partial but meaningful vulnerability information in frozen, general-purpose LLM activations and motivating model-native screening approaches.
Why it matters Shows actionable signal for vulnerability detection resides in prefill activations, suggesting low-overhead, model-native monitors as an alternative or complement to post-hoc judges/classifiers.
-
Do Time-Series Foundation Models Pay Off for Industrial Monitoring? A Cost-Aware Empirical Study
This study evaluates TSFMs versus lightweight baselines for industrial monitoring across three settings: C-MAPSS degradation risk, MIMII anomalous-sound detection with normal-only training, and BDG2 residual-diagnostics with synthetic perturbations. TCN-AE and OCSVM outperform MOMENT reconstruction on C-MAPSS and MIMII respectively, while TimesFM 2.5 yields the lowest forecast error and highest synthetic AUROC on the BDG2 panel with similar AUPRC to fitted residual models; MOMENT shows higher latency, VRAM, and state size than TCN-AE. Overall, under frozen zero-shot use, TSFMs provide task-dependent benefits rather than universally replacing compact fitted models.
Why it matters It offers a cost-aware, protocol-matched comparison showing when zero-shot TSFMs beat fitted light baselines in industrial anomaly/diagnostics workflows, informing deployment choices for secure, resource-constrained monitoring systems.
-
LabelFusion-TS: Fusing Large Language Models, Transformer Encoders, and Financial Time Series for Monetary-Policy Stance Classification
The paper introduces LabelFusion-TS, which combines a fine-tuned RoBERTa encoder, a prompted LLM, and an ensemble of time-series transformers over pre-publication market series via a small voting network to classify FOMC sentences as hawkish, dovish, or neutral. Due to limited human labels (~1k sentences), the RoBERTa encoder is first pre-trained on LLM-generated weak labels and then fine-tuned on human annotations. Trained on data up to 2015 and evaluated on 2015–2022, the fused system attains 70.2% weighted F1, surpassing a zero-shot LLM at 64.1% and overtaking it with as few as 240 human-labeled sentences.
Why it matters It demonstrates measurable gains from fusing market time series with LLM/text encoders for monetary-policy stance classification, highlighting a practical multimodal path beyond text-only classifiers.
-
Into the ORBIT for Time Series: Training Regimes for Foundation Models
The paper introduces ORBIT, a training paradigm for TSFMs that explicitly controls pre-training distributions via Bootstrap Multi-Level Sampling and Omni-Range Incremental Training. It trains Falcon-2.0, a univariate encoder-only Transformer with missingness-aware triple-channel patch tokenization and parallel patch prediction, under this regime. The authors also propose Rank-Guided Cross-Depth Alignment to align shallow and deep representations, and report strong zero-shot forecasting on GIFT-Eval and fev-bench across diverse domains and frequencies.
Why it matters It provides a principled, controllable pretraining regime and lightweight representation alignment that could transfer to secure LLM pretraining curricula and to robust TSFM deployment across heterogeneous, imbalanced time-series corpora.
-
Temporal Leakage in Financial News NLP: A Multi-Architecture Audit with a Regime-Specific M&A Signal
The paper audits temporal leakage in financial-news direction prediction across 49,799 articles and 16 model–feature setups (TF-IDF, MiniLM, FinBERT, fine-tuned RoBERTa-large/DeBERTa-v3-large, and Llama-3/Qwen2.5 zero/few-shot and LoRA), finding random splits inflate MCC by 1.1×–6.5×, increasing with capacity and feature richness, and that end-to-end FinBERT fine-tuning further amplifies the gap (size-matched ratio 1.75×). Under near-chronological evaluation, only M&A shows a positive locked-test signal (TF-IDF MCC 0.138 train-only, 0.068 after train∪val refit; 10,000-permutation p < 10^{-3}), which does not transfer to FNSPID 2009–2020 U.S. data, suggesting localization to 2024–2025 European-tilted M&A semantics. Three role labelers indicate acquirer-tagged articles as the signal locus, and the authors argue chronological splits purge stale, predictable components, leaving a small, event-localized, lexically shallow residual and call for mandatory leakage audits in financial NLP benchmarks.
Why it matters It quantifies temporal leakage effects across classic NLP and LLM settings, isolates a regime-specific residual signal, and argues for leakage audits—directly informing robust evaluation protocols for security-relevant LLMs and time-series-oriented foundation models.
Alignment
RLHF/RLAIF, preference learning, interpretability aimed at alignment, value learning, scalable oversight.
-
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families
This study reproduces Khatri et al.’s pipeline for training lightweight MLP probes on final-layer activations to detect harmful prompts, matching the original LLaMA-3.1-8B results within 0.37 F1 points (0.2 on BeaverTails). It extends the evaluation across Gemma-4-E4B, Mistral-7B-v0.3, and Qwen2-7B on WildJailbreak, BeaverTails, and AEGIS 2.0, finding F1 within about one point of the LLaMA-3.1-8B values using the same probe architecture. It also assesses nondeterminism by varying seeds and reports that final token latent vectors were invariant to seed across architectures.
Why it matters Suggests simple latent-space probes generalize across model families with stable activations, informing scalable, low-overhead safety detection for LLMs and complementing guardrail strategies.
-
ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions
ProbGuard reframes LLM safety assessment as probabilistic risk estimation over the model’s early output distribution rather than deterministic classification on completed token sequences. It estimates the probability of unsafe continuation via Monte Carlo over generated prefix distributions and calibrates this risk through post-training, enabling early stopping of unsafe generations. Empirically, it achieves the best calibration across nine model–dataset settings (average Brier and ECE reduced by 79.6% and 71.9% vs. the best baseline) and caps jailbreak attack success to ≤1% after only the first ten decoding steps.
Why it matters It operationalizes distribution-aware, early-stage safety control with strong calibration—useful for red-teaming defenses, risk-sensitive deployment, and aligning with probabilistic guardrail research and time-series-style forecasting over generation trajectories.
Agents
LLM agents: tool use, planning, multi-agent systems, agentic benchmarks and failure modes.
-
From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts
This pre-registered reproducibility audit examines LLM/agent-driven vulnerability validation artifacts across 104 papers (2023–2026), finding only 59 with publicly reachable artifacts and executing 18 paper-level workflows and 102 anchor benchmark cases. It reports high mismatch and fragility: 58/102 anchor cases have internal CVE IDs diverging from directory labels; only 10/18 artifacts run end-to-end at R0 (11/18 after R1 env-only fixes); and embedded oracles are unreliable, with 20/30 patched-counterfactuals still signaling and 7/19 negative controls triggering, yielding oracle sensitivity 60% and specificity 45%. The study argues that vulnerable-build triggers are not CVE-specific evidence without a clean patched counterfactual and offers a reusable, pre-registered protocol (post-conditions, R0/R1 repair ladder, G1–G3 evidence levels, patched-counterfactual oracles) for the community.
Why it matters It surfaces concrete failure modes and a reusable protocol for verifying LLM/agent security claims—directly informing eval design, oracle construction, and counterfactual controls for both LLM security and time-series-style benchmark reproducibility.
-
ElasticBack: Stealthy Conditional Backdoor in LLM-Agent Skills via Coupled Trigger-Rule Optimization
The paper introduces ElasticBack, a conditional single-skill backdoor for LLM agents that couples a malicious rule R in the skill document with a benign-looking trigger T in user queries, activating only when both appear. It constructs R via semantic-anchored rule injection and then fixes R while evolving T using a stealth-constrained genetic search to optimize attack effectiveness and stealth without modifying model weights. Experiments over three target behaviors (50 skills each) and four agent LLMs show high ASR with near-zero FPR, preserved clean accuracy, cross-model transfer, and evasion of deployment-time defenses, motivating stronger supply-chain defenses for skills.
Why it matters Highlights a practical, weight-free, conditional backdoor vector in the emerging agent skill ecosystem, relevant to securing tool-using LLMs and evaluating defenses against stealthy trigger-rule couplings.
-
Deployable Per-Instance Multi-Layer Activation Steering for Large Language Models
The paper shows that in activation steering, the optimal injection layers vary per input, with per-instance multi-layer selection outperforming any fixed global layer set across six binary persona traits on two 8B models. A greedy layer-ranking rule based on single-layer marginal effects nearly matches a per-instance oracle but requires gold labels, so the authors train a prompt-only predictor to mimic it and propose a deployable recipe: a prompt-based per-instance layer ranker, a classifier to infer steering direction, and an adaptive gate that limits layers by short steered passes. This deployable approach recovers most of the oracle’s gains, avoids dropping below unsteered baselines on average, reduces fluency collapse relative to strong global selection, and is supported by a mechanistic account emphasizing direction over magnitude that explains misdirection flips, collapse from over-steering, and unsteerable ceilings.
Why it matters It operationalizes per-instance, label-free activation steering with adaptive multi-layer control, offering a practical path to safer, targeted behavior edits without sacrificing fluency—relevant to red-teaming defenses and controllable generation, and conceptually aligned with direction-based causal hypotheses familiar in interpretability.
-
SkillsMetric: Mapping the Detection Boundary of Static Analysis for Malicious Agent Skills
This work introduces SkillsMetric, a five-stage static analysis framework that scores LLM agent skills across pattern density, statistical anomaly, dataflow taint, import anomaly, and capability mismatch. Using an adversarial dataset of 2,266 skills covering 16 attack types and the SkillMD-138K corpus, it reports AUC 0.93 and 5-fold CV F1 73.4%±0.5%, with 93% detection for data exfiltration and steganographic payloads. It also exposes blind spots—0% detection for host destruction via common shell commands and 42% for prompt injection—arguing static analysis alone is insufficient and should be paired with semantic review in defense-in-depth.
Why it matters Directly maps the detection boundary of static analysis for agent skills, highlighting failure modes (e.g., prompt injection, host destruction) that inform red-teaming, hybrid detection pipelines, and security benchmarks for LLM agents.
-
Generating Attacks for LLMs with GFlowNets
The work proposes an automated, human-independent red teaming framework that trains an attacker LLM via GFlowNets to probe a specified victim LLM and yield a quantitative robustness score. It addresses limitations of manual expert red teaming and automated methods tied to fixed datasets by enabling adaptive, creative attack generation. The approach targets generating more effective English adversarial inputs than existing benchmarks and, uniquely, introduces Turkish-language attack generation.
Why it matters It explores GFlowNets for adaptive attack synthesis against LLMs, offering a path beyond fixed corpora and extending multilingual red teaming—a relevant direction for security evaluation and scalable adversarial data generation.
-
Proving the Utility of Large Language Models in Cybersecurity Simulations: A Comprehensive Examination
The paper evaluates LLMs for cybersecurity simulations, using YAML to represent complex network configurations and to drive pipelines that support RL agent training. It compares LLM-based techniques to classical methods like Double Q-learning with PER, claiming improved efficiency, adaptability, and realism in cyberattack simulations. In benchmarks on multiple synthetic topologies, LLM-instantiated Python agents reached up to a 94.5% compromise rate and 0.02–0.06s per assessment, a ~25,000x–50,000x speedup over traditional RL cycles.
Why it matters Shows LLM-driven generation/execution pipelines can massively accelerate cyber-sim training and evaluation, relevant to secure LLM autonomy and sim-to-real robustness for time-series/agentic models.
-
FedLNS: Leverage LayerNorm Signature Modeling to Mitigate Adversarial Manipulation in Federated LLMs
The paper introduces FedLNS, a server-side framework that screens federated LLM client updates by modeling changes in trainable normalization-layer parameters as signatures and comparing them to a history-aware cross-client reference. It operates without additional client-to-server metadata, raw client data, trusted server datasets, labeled attack examples, or a separately trained detector, and retains full-model updates for standard or compatible aggregation after screening. Experiments on GPT-, BERT-, and LLaMA-style models trained from scratch with 200 clients show that under 40% target manipulation, FedLNS yields lower test perplexity than six baselines across IID and non-IID partitions.
Why it matters Offers a practical, data-free, model-internal signal for adversarial update filtering in federated LLM training, aligning with robustness and integrity concerns in LLM security.
-
RouteGuard: Certifying Routing Gain in LLM Multi-Agent Systems When Complementarity Is Not Enough
The paper shows that common assumptions for multi-agent LLM routing—optimizing gate AUC and relying on advisor complementarity—do not determine deployable gain, and introduces RouteGuard to certify routing benefit. It decomposes gain as G = πΔ_E, identifies a conditional-regret functional Φ as the governing quantity (not AUC), provides a finite-sample certification bracket with a matching Le Cam lower bound that is constant-sharp over the fixed-activity class, and observes a robustness phase transition. Experiments on RouterBench and OpenRCA demonstrate RouteGuard acting as a guardrail, certifying gain only under prompt-level sampling (not workload-cluster resampling) on RouterBench and refusing to certify on OpenRCA due to advisor redundancy, with pre-registered semi-synthetic controls confirming calibration thresholds.
Why it matters It offers a principled, sample-sensitive certification of routing gain that can prevent illusory improvements in LLM multi-agent systems and aligns evaluation with deployment risk.
-
Stopping and Routing LLM Judge Panels
The work frames judge-panel design as a role-conditioned allocation problem that, using a small labeled audit set, declared slices, and judge costs, estimates target-relative roles: copies, complements, and specialists. These roles yield a policy to drop copies, add complements globally, route specialists on slices, and stop when validation gain falls below a threshold, producing a reusable, auditable call plan. It is empirically compared across multiple audit settings against single judges, flat panels, diversity heuristics, full-call stacking, reliability juries, and frugal cascades, yielding a regime map for when to route specialists, stop in saturated regimes, keep broad ensembles, and ignore conditional copies.
Why it matters It offers a concrete, auditable routing/early-stopping policy for multi-judge LLM evaluation that balances risk and cost across slices—directly relevant to secure deployment and panel design in LLM oversight and verifiers for time-series or reasoning tasks.
-
LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications
The paper surveys LLM-based forecasting agents, organizing architectures into standalone LLM workflows over temporal inputs, tool/retrieval-augmented agents, and hybrids with statistical or foundation models. It reviews training and evaluation practices and highlights negative evidence such as sensitivity to small perturbations, ablations showing no accuracy gains from the LLM, and potential benchmark contamination confounding temporal reasoning claims. Applications span finance, weather, health, energy, and operations, with the core limitation identified as measurement, motivating needs for robust calibration under shift, contamination-resistant live evaluation, explicit cost–accuracy reporting, and methods addressing forecast-induced feedback effects.
Why it matters Concise map of architectures, pitfalls, and eval gaps for LLM-in-the-loop forecasting—useful for security-minded eval design, contamination control, calibration under shift, and coupling with time-series FMs.
-
FutureBridge: Token Selection Beyond Local Preference in Collaborative Decoding
FutureBridge proposes a token reranking scheme for LLM–SLM collaborative decoding that prioritizes tokens enabling the SLM’s subsequent reasoning rather than the LLM’s local preference. Training uses an answer-verified LLM trajectory to fix a shared future, while a frozen SLM assigns counterfactual scores to candidate tokens under this common context to supervise a lightweight reranker observing only the current state and token. At inference, the LLM only expands the candidate pool and the reranker selects a single token before handing generation back to the SLM without appending any future suffix, yielding a 35.1% relative Math Avg. improvement for Qwen3-1.7B across five math reasoning benchmarks over greedy decoding.
Why it matters It reframes collaboration as SLM-aware token selection, offering a practical path to safer, stronger small-model reasoning with minimal LLM involvement—relevant to alignment, tool-use gating, and cost/security tradeoffs in cooperative decoding.
-
Do Time-Series Foundation Models Pay Off for Industrial Monitoring? A Cost-Aware Empirical Study
This study evaluates TSFMs versus lightweight baselines for industrial monitoring across three settings: C-MAPSS degradation risk, MIMII anomalous-sound detection with normal-only training, and BDG2 residual-diagnostics with synthetic perturbations. TCN-AE and OCSVM outperform MOMENT reconstruction on C-MAPSS and MIMII respectively, while TimesFM 2.5 yields the lowest forecast error and highest synthetic AUROC on the BDG2 panel with similar AUPRC to fitted residual models; MOMENT shows higher latency, VRAM, and state size than TCN-AE. Overall, under frozen zero-shot use, TSFMs provide task-dependent benefits rather than universally replacing compact fitted models.
Why it matters It offers a cost-aware, protocol-matched comparison showing when zero-shot TSFMs beat fitted light baselines in industrial anomaly/diagnostics workflows, informing deployment choices for secure, resource-constrained monitoring systems.
-
DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model
DreamGuard is a proactive runtime guardrail for LLM agents that uses a risk-aware world model to track a compact recurrent latent state over trajectories and forecast future latent states. From these forecasts, it derives immediate-hazard and prefix-risk evidence and fuses multi-horizon signals to decide on interventions before tool executions. Across four benchmarks and an online evaluation, DreamGuard outperforms generic, reactive, and proactive baselines with the best safety-utility trade-off and averages 25 ms end-to-end latency per call.
Why it matters Introduces a trajectory-aware, low-latency guardrail with explicit risk modeling, aligning with interests in long-horizon safety for LLM agents and bridging to sequence modeling ideas from time series foundation models.
-
Learning the Pareto Frontier of Predictive Models under Distribution Shift
The paper introduces Frontier Learning, which treats a library of pretrained models with mixed access regimes (black-box predictions and white-box representations) as complementary sources under distribution shift. It builds a unified target-domain feature by concatenating internal representations from white-box models with outputs from black-box models, then trains a lightweight regularized supervised learner, guaranteeing training-sample risk no worse than any single baseline including zero-shot reuse, fine-tuning, or direct training. Experiments on simulations and real-world shifts (DomainNet/VisDA and MIMIC-IV-Notes ICU mortality) show it matches or outperforms the strongest individual reuse strategy, with the largest gains when no single baseline is reliable across the shifts considered.
Why it matters Provides a simple, access-agnostic ensemble-on-representations approach that subsumes common reuse strategies and is empirically strong under shift—relevant to robustness/security and modular reuse of foundation models.
-
ThreatLens: Evidence-Guided Ranking of High-Priority CVEs
ThreatLens is a deployment-realistic CVE prioritization framework that ranks at review time using only cutoff-valid evidence and learns exploitation relevance from future CISA KEV entries as weak supervision. Under forward-in-time, CVE-disjoint evaluation, it significantly outperforms CVSS, EPSS, and rule-based evidence fusion, surfacing 80.0% of future KEV CVEs in the top 20 and 95.9% in the top 50 on a held-out test split. Early-warning analysis shows it flags a substantial fraction of later KEV entries before catalog inclusion, enabling timely, evidence-grounded triage.
Why it matters Demonstrates a realistic, temporally valid ranking approach with weak supervision that materially beats common baselines—relevant to LLM security triage pipelines and time-aware evaluation of predictive models.
-
Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning
Palmyra x6 is an enterprise-focused agentic LLM post-trained from an MoE base using Anchored Supervised Fine-Tuning on 626 verified synthetic tool-use trajectories, run for a single epoch with a low LR and a KL anchor to the frozen base, optimized via a Muon+Adam hybrid. It delivers substantial gains for the Writer Agent and compares favorably on public benchmarks, achieving the top BFCL Core score of 0.785 and the highest six-benchmark mean among its cohort. The model is also competitive or leading in the authors’ bias and safety evaluations relative to comparators.
Why it matters Tight, KL-anchored post-training on compact synthetic tool-use traces yielding SOTA-ish agentic performance is directly relevant to secure tool-use alignment and data-efficient post-training strategies for robust, controllable agents.
Efficiency
Inference/training efficiency: quantization, distillation, sparsity, serving systems, hardware-aware methods.
-
Deployable Per-Instance Multi-Layer Activation Steering for Large Language Models
The paper shows that in activation steering, the optimal injection layers vary per input, with per-instance multi-layer selection outperforming any fixed global layer set across six binary persona traits on two 8B models. A greedy layer-ranking rule based on single-layer marginal effects nearly matches a per-instance oracle but requires gold labels, so the authors train a prompt-only predictor to mimic it and propose a deployable recipe: a prompt-based per-instance layer ranker, a classifier to infer steering direction, and an adaptive gate that limits layers by short steered passes. This deployable approach recovers most of the oracle’s gains, avoids dropping below unsteered baselines on average, reduces fluency collapse relative to strong global selection, and is supported by a mechanistic account emphasizing direction over magnitude that explains misdirection flips, collapse from over-steering, and unsteerable ceilings.
Why it matters It operationalizes per-instance, label-free activation steering with adaptive multi-layer control, offering a practical path to safer, targeted behavior edits without sacrificing fluency—relevant to red-teaming defenses and controllable generation, and conceptually aligned with direction-based causal hypotheses familiar in interpretability.
-
FutureBridge: Token Selection Beyond Local Preference in Collaborative Decoding
FutureBridge proposes a token reranking scheme for LLM–SLM collaborative decoding that prioritizes tokens enabling the SLM’s subsequent reasoning rather than the LLM’s local preference. Training uses an answer-verified LLM trajectory to fix a shared future, while a frozen SLM assigns counterfactual scores to candidate tokens under this common context to supervise a lightweight reranker observing only the current state and token. At inference, the LLM only expands the candidate pool and the reranker selects a single token before handing generation back to the SLM without appending any future suffix, yielding a 35.1% relative Math Avg. improvement for Qwen3-1.7B across five math reasoning benchmarks over greedy decoding.
Why it matters It reframes collaboration as SLM-aware token selection, offering a practical path to safer, stronger small-model reasoning with minimal LLM involvement—relevant to alignment, tool-use gating, and cost/security tradeoffs in cooperative decoding.
-
DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model
DreamGuard is a proactive runtime guardrail for LLM agents that uses a risk-aware world model to track a compact recurrent latent state over trajectories and forecast future latent states. From these forecasts, it derives immediate-hazard and prefix-risk evidence and fuses multi-horizon signals to decide on interventions before tool executions. Across four benchmarks and an online evaluation, DreamGuard outperforms generic, reactive, and proactive baselines with the best safety-utility trade-off and averages 25 ms end-to-end latency per call.
Why it matters Introduces a trajectory-aware, low-latency guardrail with explicit risk modeling, aligning with interests in long-horizon safety for LLM agents and bridging to sequence modeling ideas from time series foundation models.
-
Learning the Pareto Frontier of Predictive Models under Distribution Shift
The paper introduces Frontier Learning, which treats a library of pretrained models with mixed access regimes (black-box predictions and white-box representations) as complementary sources under distribution shift. It builds a unified target-domain feature by concatenating internal representations from white-box models with outputs from black-box models, then trains a lightweight regularized supervised learner, guaranteeing training-sample risk no worse than any single baseline including zero-shot reuse, fine-tuning, or direct training. Experiments on simulations and real-world shifts (DomainNet/VisDA and MIMIC-IV-Notes ICU mortality) show it matches or outperforms the strongest individual reuse strategy, with the largest gains when no single baseline is reliable across the shifts considered.
Why it matters Provides a simple, access-agnostic ensemble-on-representations approach that subsumes common reuse strategies and is empirically strong under shift—relevant to robustness/security and modular reuse of foundation models.
Evaluation
Benchmarks, evaluation methodology, dataset/leakage critique, LLM-as-judge.
-
From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts
This pre-registered reproducibility audit examines LLM/agent-driven vulnerability validation artifacts across 104 papers (2023–2026), finding only 59 with publicly reachable artifacts and executing 18 paper-level workflows and 102 anchor benchmark cases. It reports high mismatch and fragility: 58/102 anchor cases have internal CVE IDs diverging from directory labels; only 10/18 artifacts run end-to-end at R0 (11/18 after R1 env-only fixes); and embedded oracles are unreliable, with 20/30 patched-counterfactuals still signaling and 7/19 negative controls triggering, yielding oracle sensitivity 60% and specificity 45%. The study argues that vulnerable-build triggers are not CVE-specific evidence without a clean patched counterfactual and offers a reusable, pre-registered protocol (post-conditions, R0/R1 repair ladder, G1–G3 evidence levels, patched-counterfactual oracles) for the community.
Why it matters It surfaces concrete failure modes and a reusable protocol for verifying LLM/agent security claims—directly informing eval design, oracle construction, and counterfactual controls for both LLM security and time-series-style benchmark reproducibility.
-
ElasticBack: Stealthy Conditional Backdoor in LLM-Agent Skills via Coupled Trigger-Rule Optimization
The paper introduces ElasticBack, a conditional single-skill backdoor for LLM agents that couples a malicious rule R in the skill document with a benign-looking trigger T in user queries, activating only when both appear. It constructs R via semantic-anchored rule injection and then fixes R while evolving T using a stealth-constrained genetic search to optimize attack effectiveness and stealth without modifying model weights. Experiments over three target behaviors (50 skills each) and four agent LLMs show high ASR with near-zero FPR, preserved clean accuracy, cross-model transfer, and evasion of deployment-time defenses, motivating stronger supply-chain defenses for skills.
Why it matters Highlights a practical, weight-free, conditional backdoor vector in the emerging agent skill ecosystem, relevant to securing tool-using LLMs and evaluating defenses against stealthy trigger-rule couplings.
-
Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection
This work evaluates indirect prompt injection against DeepSeek Harness using AI-Infra-Guard to generate tests, deliver controlled taint across 16 channels and two carrier modes, run 14,560 executions, and analyze traces with both rule-based and LLM-based judges. The strongest observed attack success rates are 17.0% under LLMJudge for a fake-completion attack (text), 25.5% under RuleJudge for hidden Unicode (file), and 16.0% under RuleJudge for the skills channel (file), with LLMJudge assigning more partial compliance than RuleJudge (7.3% vs. 2.0%). The study connects outcomes to DSH’s handling of tool results, extra contexts, and tool-call policy hooks, and proposes controls between untrusted content and sensitive actions, with code released at the linked repository.
Why it matters Provides a large-scale, reproducible, channel- and payload-stratified benchmark of indirect prompt injection on an agentic LLM stack with dual judges, informing defense design and evaluation baselines.
-
LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models
The paper introduces LiveHouse-TS, an open-world, living benchmark infrastructure for evaluating TSFMs via prequential testing on real future data rather than static windows. It shifts evaluation focus from snapshot accuracy to continuous temporal validity in environments with seasonality, distribution shifts, and unexpected events, aiming to answer long-term questions about ranking stability and robustness. Streaming evaluations across 11 domains and 17 datasets show that model rankings from static benchmarks dramatically reshuffle under the live protocol.
Why it matters If you work on LLM security or TSFMs, this exposes how static benchmarks can mislead robustness assessments under shifts, offering an infrastructure to stress-test models’ long-term validity—analogous to live red-teaming and continual evals in LLMs.
-
When Do PEFT Adaptations Leak Structure? Measuring Black-Box Structural Bounds in Public-Base Model Services
The paper introduces VectorHijack-SR, a black-box measurement method that turns paired victim/base residuals into calibrated bounds over PEFT family, layer locality, and coarse rank, using aggregated query-level statistics and a service-disjoint classifier plus a cross-fitted hierarchical rejector to test LoRA-manifold membership. Empirically, family leakage exceeds chance across multiple backbones/tasks, rank inference varies by task, the rejector attains AUROC 0.804 and high known-set accuracy but struggles with structurally close DoRA/LoRA+head, and exact-version linkage on held-out LoRA-r64 services reaches AUC 0.940. Despite measurable structural leakage, experiments show a visibility–exploitability gap: two-stage recovery offers no fair-budget query savings, posterior-selected PEFT underperforms distill-then-convert PEFT, and free-running generation remains near chance.
Why it matters It provides concrete, black-box evidence and metrics for structural leakage from PEFTed services with known bases, informing red-team audits, model fingerprinting, and security evaluations of adapter deployment choices.
-
Generating Attacks for LLMs with GFlowNets
The work proposes an automated, human-independent red teaming framework that trains an attacker LLM via GFlowNets to probe a specified victim LLM and yield a quantitative robustness score. It addresses limitations of manual expert red teaming and automated methods tied to fixed datasets by enabling adaptive, creative attack generation. The approach targets generating more effective English adversarial inputs than existing benchmarks and, uniquely, introduces Turkish-language attack generation.
Why it matters It explores GFlowNets for adaptive attack synthesis against LLMs, offering a path beyond fixed corpora and extending multilingual red teaming—a relevant direction for security evaluation and scalable adversarial data generation.
-
Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining
The paper introduces a task-agnostic measure of training data influence that quantifies how much an example’s gradient update reduces the squared distance to a run’s final parameters, estimated from intermediate checkpoints without retraining. Applied to 18 Pythia and PolyPythia configurations, the method reveals systematic temporal shifts in influential data: literature-related data align more with the parameter trajectory early, while STEM data align more later. This qualitative crossover is broadly consistent across configurations, offering a tractable trajectory-level perspective that complements task-based influence analyses.
Why it matters Gives a practical, retraining-free way to map which data sources steer pretraining trajectories over time, informing data curation, curriculum, and security-relevant provenance audits for LLMs.
-
Proving the Utility of Large Language Models in Cybersecurity Simulations: A Comprehensive Examination
The paper evaluates LLMs for cybersecurity simulations, using YAML to represent complex network configurations and to drive pipelines that support RL agent training. It compares LLM-based techniques to classical methods like Double Q-learning with PER, claiming improved efficiency, adaptability, and realism in cyberattack simulations. In benchmarks on multiple synthetic topologies, LLM-instantiated Python agents reached up to a 94.5% compromise rate and 0.02–0.06s per assessment, a ~25,000x–50,000x speedup over traditional RL cycles.
Why it matters Shows LLM-driven generation/execution pipelines can massively accelerate cyber-sim training and evaluation, relevant to secure LLM autonomy and sim-to-real robustness for time-series/agentic models.
-
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families
This study reproduces Khatri et al.’s pipeline for training lightweight MLP probes on final-layer activations to detect harmful prompts, matching the original LLaMA-3.1-8B results within 0.37 F1 points (0.2 on BeaverTails). It extends the evaluation across Gemma-4-E4B, Mistral-7B-v0.3, and Qwen2-7B on WildJailbreak, BeaverTails, and AEGIS 2.0, finding F1 within about one point of the LLaMA-3.1-8B values using the same probe architecture. It also assesses nondeterminism by varying seeds and reports that final token latent vectors were invariant to seed across architectures.
Why it matters Suggests simple latent-space probes generalize across model families with stable activations, informing scalable, low-overhead safety detection for LLMs and complementing guardrail strategies.
-
RouteGuard: Certifying Routing Gain in LLM Multi-Agent Systems When Complementarity Is Not Enough
The paper shows that common assumptions for multi-agent LLM routing—optimizing gate AUC and relying on advisor complementarity—do not determine deployable gain, and introduces RouteGuard to certify routing benefit. It decomposes gain as G = πΔ_E, identifies a conditional-regret functional Φ as the governing quantity (not AUC), provides a finite-sample certification bracket with a matching Le Cam lower bound that is constant-sharp over the fixed-activity class, and observes a robustness phase transition. Experiments on RouterBench and OpenRCA demonstrate RouteGuard acting as a guardrail, certifying gain only under prompt-level sampling (not workload-cluster resampling) on RouterBench and refusing to certify on OpenRCA due to advisor redundancy, with pre-registered semi-synthetic controls confirming calibration thresholds.
Why it matters It offers a principled, sample-sensitive certification of routing gain that can prevent illusory improvements in LLM multi-agent systems and aligns evaluation with deployment risk.
-
Measure, Don't Optimize: Forecasting Recovery in LLM Unlearning
The paper introduces J-Access, an inference-time audit that projects intermediate representations into vocabulary space to measure how often target concepts remain accessible along the output pathway, hypothesizing that residual accessibility predicts recovery susceptibility. Auditing 398 public unlearned models across eight methods, they find most retain accessibility above a retain-only gold level, pre-attack accessibility predicts recovery speed and extent at the model level (but not per-fact), and directly minimizing J-Access causes the model to hide knowledge from the audit while increasing post-attack recovery. They conclude J-Access is useful as a model-level diagnostic for residual susceptibility and caution against turning internal audits into optimization targets without validation.
Why it matters Offers an actionable, model-level risk signal for unlearning robustness and a clear warning about Goodharting internal audits—relevant for designing secure unlearning and audit pipelines in LLMs.
-
Stopping and Routing LLM Judge Panels
The work frames judge-panel design as a role-conditioned allocation problem that, using a small labeled audit set, declared slices, and judge costs, estimates target-relative roles: copies, complements, and specialists. These roles yield a policy to drop copies, add complements globally, route specialists on slices, and stop when validation gain falls below a threshold, producing a reusable, auditable call plan. It is empirically compared across multiple audit settings against single judges, flat panels, diversity heuristics, full-call stacking, reliability juries, and frugal cascades, yielding a regime map for when to route specialists, stop in saturated regimes, keep broad ensembles, and ignore conditional copies.
Why it matters It offers a concrete, auditable routing/early-stopping policy for multi-judge LLM evaluation that balances risk and cost across slices—directly relevant to secure deployment and panel design in LLM oversight and verifiers for time-series or reasoning tasks.
-
LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications
The paper surveys LLM-based forecasting agents, organizing architectures into standalone LLM workflows over temporal inputs, tool/retrieval-augmented agents, and hybrids with statistical or foundation models. It reviews training and evaluation practices and highlights negative evidence such as sensitivity to small perturbations, ablations showing no accuracy gains from the LLM, and potential benchmark contamination confounding temporal reasoning claims. Applications span finance, weather, health, energy, and operations, with the core limitation identified as measurement, motivating needs for robust calibration under shift, contamination-resistant live evaluation, explicit cost–accuracy reporting, and methods addressing forecast-induced feedback effects.
Why it matters Concise map of architectures, pitfalls, and eval gaps for LLM-in-the-loop forecasting—useful for security-minded eval design, contamination control, calibration under shift, and coupling with time-series FMs.
-
ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions
ProbGuard reframes LLM safety assessment as probabilistic risk estimation over the model’s early output distribution rather than deterministic classification on completed token sequences. It estimates the probability of unsafe continuation via Monte Carlo over generated prefix distributions and calibrates this risk through post-training, enabling early stopping of unsafe generations. Empirically, it achieves the best calibration across nine model–dataset settings (average Brier and ECE reduced by 79.6% and 71.9% vs. the best baseline) and caps jailbreak attack success to ≤1% after only the first ten decoding steps.
Why it matters It operationalizes distribution-aware, early-stage safety control with strong calibration—useful for red-teaming defenses, risk-sensitive deployment, and aligning with probabilistic guardrail research and time-series-style forecasting over generation trajectories.
-
Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations
The authors test whether LLM hidden states during prefill encode vulnerability signals when reading C/C++ code by extracting last prefill token activations from four models and training small MLP probes. Across four function-level benchmarks (Devign, Big-Vul, Draper VDISC, PrimeVul), probes (13.4–16.0M params, <0.2% of base) reach 41.7% average F1, with a Qwen3.5-9B probe hitting 68.8% F1 on Devign, comparable to a fine-tuned-classifier SOTA (67.9%). Performance lags SOTA on the more imbalanced datasets, indicating partial but meaningful vulnerability information in frozen, general-purpose LLM activations and motivating model-native screening approaches.
Why it matters Shows actionable signal for vulnerability detection resides in prefill activations, suggesting low-overhead, model-native monitors as an alternative or complement to post-hoc judges/classifiers.
-
Do Time-Series Foundation Models Pay Off for Industrial Monitoring? A Cost-Aware Empirical Study
This study evaluates TSFMs versus lightweight baselines for industrial monitoring across three settings: C-MAPSS degradation risk, MIMII anomalous-sound detection with normal-only training, and BDG2 residual-diagnostics with synthetic perturbations. TCN-AE and OCSVM outperform MOMENT reconstruction on C-MAPSS and MIMII respectively, while TimesFM 2.5 yields the lowest forecast error and highest synthetic AUROC on the BDG2 panel with similar AUPRC to fitted residual models; MOMENT shows higher latency, VRAM, and state size than TCN-AE. Overall, under frozen zero-shot use, TSFMs provide task-dependent benefits rather than universally replacing compact fitted models.
Why it matters It offers a cost-aware, protocol-matched comparison showing when zero-shot TSFMs beat fitted light baselines in industrial anomaly/diagnostics workflows, informing deployment choices for secure, resource-constrained monitoring systems.
-
DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model
DreamGuard is a proactive runtime guardrail for LLM agents that uses a risk-aware world model to track a compact recurrent latent state over trajectories and forecast future latent states. From these forecasts, it derives immediate-hazard and prefix-risk evidence and fuses multi-horizon signals to decide on interventions before tool executions. Across four benchmarks and an online evaluation, DreamGuard outperforms generic, reactive, and proactive baselines with the best safety-utility trade-off and averages 25 ms end-to-end latency per call.
Why it matters Introduces a trajectory-aware, low-latency guardrail with explicit risk modeling, aligning with interests in long-horizon safety for LLM agents and bridging to sequence modeling ideas from time series foundation models.
-
LabelFusion-TS: Fusing Large Language Models, Transformer Encoders, and Financial Time Series for Monetary-Policy Stance Classification
The paper introduces LabelFusion-TS, which combines a fine-tuned RoBERTa encoder, a prompted LLM, and an ensemble of time-series transformers over pre-publication market series via a small voting network to classify FOMC sentences as hawkish, dovish, or neutral. Due to limited human labels (~1k sentences), the RoBERTa encoder is first pre-trained on LLM-generated weak labels and then fine-tuned on human annotations. Trained on data up to 2015 and evaluated on 2015–2022, the fused system attains 70.2% weighted F1, surpassing a zero-shot LLM at 64.1% and overtaking it with as few as 240 human-labeled sentences.
Why it matters It demonstrates measurable gains from fusing market time series with LLM/text encoders for monetary-policy stance classification, highlighting a practical multimodal path beyond text-only classifiers.
-
Analysis of Federated Aggregation under Model Poisoning and Backdoor Attacks: A Reconstructed Cross-Dataset and Cross-Architecture Benchmark
This work reconstructs a 500-cell federated aggregation benchmark across five methods, five datasets, five architectures, and four conditions (clean, sign-flipping, Gaussian, BadNets), recovering 454 original runs and 36 repairs/reruns, with 10 SVHN cells covered only by summary provenance. Trimmed Mean has the best clean macro-mean accuracy and lowest mean within-task rank, while Krum shows the top recorded accuracy under sign-flipping and Gaussian, and these rankings persist when restricted to fully logged task pairs. Audits reveal the BadNets metric is actually TTLR (all test inputs triggered) and a FedPARETO pathway where reported predictive summaries may not match the corrupted updates used for aggregation, so results are descriptive within recorded configs rather than universal robustness claims.
Why it matters Highlights reproducibility gaps, metric mis-specification (TTLR), and aggregation/provenance pitfalls that can confound robustness claims in federated settings relevant to secure FL and poisoning/backdoor evaluation.
-
Into the ORBIT for Time Series: Training Regimes for Foundation Models
The paper introduces ORBIT, a training paradigm for TSFMs that explicitly controls pre-training distributions via Bootstrap Multi-Level Sampling and Omni-Range Incremental Training. It trains Falcon-2.0, a univariate encoder-only Transformer with missingness-aware triple-channel patch tokenization and parallel patch prediction, under this regime. The authors also propose Rank-Guided Cross-Depth Alignment to align shallow and deep representations, and report strong zero-shot forecasting on GIFT-Eval and fev-bench across diverse domains and frequencies.
Why it matters It provides a principled, controllable pretraining regime and lightweight representation alignment that could transfer to secure LLM pretraining curricula and to robust TSFM deployment across heterogeneous, imbalanced time-series corpora.
-
ThreatLens: Evidence-Guided Ranking of High-Priority CVEs
ThreatLens is a deployment-realistic CVE prioritization framework that ranks at review time using only cutoff-valid evidence and learns exploitation relevance from future CISA KEV entries as weak supervision. Under forward-in-time, CVE-disjoint evaluation, it significantly outperforms CVSS, EPSS, and rule-based evidence fusion, surfacing 80.0% of future KEV CVEs in the top 20 and 95.9% in the top 50 on a held-out test split. Early-warning analysis shows it flags a substantial fraction of later KEV entries before catalog inclusion, enabling timely, evidence-grounded triage.
Why it matters Demonstrates a realistic, temporally valid ranking approach with weak supervision that materially beats common baselines—relevant to LLM security triage pipelines and time-aware evaluation of predictive models.
-
Cross-Corpus Evaluation of Generalizable Vulnerability Detection in IoT Firmware
The paper introduces IoTVulBench, a human-verified, contamination-screened benchmark for cross-corpus IoT firmware vulnerability detection, and evaluates five model architectures, two tuning methods, and three curriculum strategies with ensemble, distillation, and robustness analyses. Models trained on IoTVulBench outperform matched single-source datasets (MCC 0.58 vs. 0.44 for PrimeVul and 0.39 for D2A), with staged curriculum learning raising MCC to 0.69 and a diversity-optimized ensemble reaching 0.73, surpassing a static analyzer (0.31) and PrimeVul by large margins while maintaining 86% performance under identifier renaming and showing strong calibration and largely faithful explanations. At a 0.5% FPR, the model missed 21% of vulnerabilities versus 71% for the strongest comparator, indicating domain-matched data and curriculum design—not model scale alone—drive generalization in firmware vulnerability detection.
Why it matters Provides a contamination-screened, human-verified benchmark and concrete training/curriculum/ensemble recipes that substantially improve cross-corpus generalization—useful for LLM security evaluation and for designing data- and curriculum-centric pipelines beyond scaling.
-
Temporal Leakage in Financial News NLP: A Multi-Architecture Audit with a Regime-Specific M&A Signal
The paper audits temporal leakage in financial-news direction prediction across 49,799 articles and 16 model–feature setups (TF-IDF, MiniLM, FinBERT, fine-tuned RoBERTa-large/DeBERTa-v3-large, and Llama-3/Qwen2.5 zero/few-shot and LoRA), finding random splits inflate MCC by 1.1×–6.5×, increasing with capacity and feature richness, and that end-to-end FinBERT fine-tuning further amplifies the gap (size-matched ratio 1.75×). Under near-chronological evaluation, only M&A shows a positive locked-test signal (TF-IDF MCC 0.138 train-only, 0.068 after train∪val refit; 10,000-permutation p < 10^{-3}), which does not transfer to FNSPID 2009–2020 U.S. data, suggesting localization to 2024–2025 European-tilted M&A semantics. Three role labelers indicate acquirer-tagged articles as the signal locus, and the authors argue chronological splits purge stale, predictable components, leaving a small, event-localized, lexically shallow residual and call for mandatory leakage audits in financial NLP benchmarks.
Why it matters It quantifies temporal leakage effects across classic NLP and LLM settings, isolates a regime-specific residual signal, and argues for leakage audits—directly informing robust evaluation protocols for security-relevant LLMs and time-series-oriented foundation models.
Ecosystem
Model releases
New model weights or major model announcements from labs.
-
Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection
This work evaluates indirect prompt injection against DeepSeek Harness using AI-Infra-Guard to generate tests, deliver controlled taint across 16 channels and two carrier modes, run 14,560 executions, and analyze traces with both rule-based and LLM-based judges. The strongest observed attack success rates are 17.0% under LLMJudge for a fake-completion attack (text), 25.5% under RuleJudge for hidden Unicode (file), and 16.0% under RuleJudge for the skills channel (file), with LLMJudge assigning more partial compliance than RuleJudge (7.3% vs. 2.0%). The study connects outcomes to DSH’s handling of tool results, extra contexts, and tool-call policy hooks, and proposes controls between untrusted content and sensitive actions, with code released at the linked repository.
Why it matters Provides a large-scale, reproducible, channel- and payload-stratified benchmark of indirect prompt injection on an agentic LLM stack with dual judges, informing defense design and evaluation baselines.
-
Measure, Don't Optimize: Forecasting Recovery in LLM Unlearning
The paper introduces J-Access, an inference-time audit that projects intermediate representations into vocabulary space to measure how often target concepts remain accessible along the output pathway, hypothesizing that residual accessibility predicts recovery susceptibility. Auditing 398 public unlearned models across eight methods, they find most retain accessibility above a retain-only gold level, pre-attack accessibility predicts recovery speed and extent at the model level (but not per-fact), and directly minimizing J-Access causes the model to hide knowledge from the audit while increasing post-attack recovery. They conclude J-Access is useful as a model-level diagnostic for residual susceptibility and caution against turning internal audits into optimization targets without validation.
Why it matters Offers an actionable, model-level risk signal for unlearning robustness and a clear warning about Goodharting internal audits—relevant for designing secure unlearning and audit pipelines in LLMs.
-
Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning
Palmyra x6 is an enterprise-focused agentic LLM post-trained from an MoE base using Anchored Supervised Fine-Tuning on 626 verified synthetic tool-use trajectories, run for a single epoch with a low LR and a KL anchor to the frozen base, optimized via a Muon+Adam hybrid. It delivers substantial gains for the Writer Agent and compares favorably on public benchmarks, achieving the top BFCL Core score of 0.785 and the highest six-benchmark mean among its cohort. The model is also competitive or leading in the authors’ bias and safety evaluations relative to comparators.
Why it matters Tight, KL-anchored post-training on compact synthetic tool-use traces yielding SOTA-ish agentic performance is directly relevant to secure tool-use alignment and data-efficient post-training strategies for robust, controllable agents.
Policy
AI policy, regulation, governance, export controls.
Longform
Essays, blog posts, and talks worth the extended read — not papers.
-
From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts
This pre-registered reproducibility audit examines LLM/agent-driven vulnerability validation artifacts across 104 papers (2023–2026), finding only 59 with publicly reachable artifacts and executing 18 paper-level workflows and 102 anchor benchmark cases. It reports high mismatch and fragility: 58/102 anchor cases have internal CVE IDs diverging from directory labels; only 10/18 artifacts run end-to-end at R0 (11/18 after R1 env-only fixes); and embedded oracles are unreliable, with 20/30 patched-counterfactuals still signaling and 7/19 negative controls triggering, yielding oracle sensitivity 60% and specificity 45%. The study argues that vulnerable-build triggers are not CVE-specific evidence without a clean patched counterfactual and offers a reusable, pre-registered protocol (post-conditions, R0/R1 repair ladder, G1–G3 evidence levels, patched-counterfactual oracles) for the community.
Why it matters It surfaces concrete failure modes and a reusable protocol for verifying LLM/agent security claims—directly informing eval design, oracle construction, and counterfactual controls for both LLM security and time-series-style benchmark reproducibility.
-
LabelFusion-TS: Fusing Large Language Models, Transformer Encoders, and Financial Time Series for Monetary-Policy Stance Classification
The paper introduces LabelFusion-TS, which combines a fine-tuned RoBERTa encoder, a prompted LLM, and an ensemble of time-series transformers over pre-publication market series via a small voting network to classify FOMC sentences as hawkish, dovish, or neutral. Due to limited human labels (~1k sentences), the RoBERTa encoder is first pre-trained on LLM-generated weak labels and then fine-tuned on human annotations. Trained on data up to 2015 and evaluated on 2015–2022, the fused system attains 70.2% weighted F1, surpassing a zero-shot LLM at 64.1% and overtaking it with as few as 240 human-labeled sentences.
Why it matters It demonstrates measurable gains from fusing market time series with LLM/text encoders for monetary-policy stance classification, highlighting a practical multimodal path beyond text-only classifiers.
-
Temporal Leakage in Financial News NLP: A Multi-Architecture Audit with a Regime-Specific M&A Signal
The paper audits temporal leakage in financial-news direction prediction across 49,799 articles and 16 model–feature setups (TF-IDF, MiniLM, FinBERT, fine-tuned RoBERTa-large/DeBERTa-v3-large, and Llama-3/Qwen2.5 zero/few-shot and LoRA), finding random splits inflate MCC by 1.1×–6.5×, increasing with capacity and feature richness, and that end-to-end FinBERT fine-tuning further amplifies the gap (size-matched ratio 1.75×). Under near-chronological evaluation, only M&A shows a positive locked-test signal (TF-IDF MCC 0.138 train-only, 0.068 after train∪val refit; 10,000-permutation p < 10^{-3}), which does not transfer to FNSPID 2009–2020 U.S. data, suggesting localization to 2024–2025 European-tilted M&A semantics. Three role labelers indicate acquirer-tagged articles as the signal locus, and the authors argue chronological splits purge stale, predictable components, leaving a small, event-localized, lexically shallow residual and call for mandatory leakage audits in financial NLP benchmarks.
Why it matters It quantifies temporal leakage effects across classic NLP and LLM settings, isolates a regime-specific residual signal, and argues for leakage audits—directly informing robust evaluation protocols for security-relevant LLMs and time-series-oriented foundation models.
Utility
Has code
Filter, not a ranker — papers in the user's areas with a working code repo linked.
-
Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection
This work evaluates indirect prompt injection against DeepSeek Harness using AI-Infra-Guard to generate tests, deliver controlled taint across 16 channels and two carrier modes, run 14,560 executions, and analyze traces with both rule-based and LLM-based judges. The strongest observed attack success rates are 17.0% under LLMJudge for a fake-completion attack (text), 25.5% under RuleJudge for hidden Unicode (file), and 16.0% under RuleJudge for the skills channel (file), with LLMJudge assigning more partial compliance than RuleJudge (7.3% vs. 2.0%). The study connects outcomes to DSH’s handling of tool results, extra contexts, and tool-call policy hooks, and proposes controls between untrusted content and sensitive actions, with code released at the linked repository.
Why it matters Provides a large-scale, reproducible, channel- and payload-stratified benchmark of indirect prompt injection on an agentic LLM stack with dual judges, informing defense design and evaluation baselines.
Black-box auditing of vendor-hosted LLM APIs via statistical repeated-query testing is highly relevant for red-teaming, API threat models, and evaluation methodology.
-
COPEX: Benchmarking LLM Robustness to Adversarial Context Across Model Context Protocol Layers
Directly targets adversarial robustness benchmarks for LLM-based MCP/security workflows, aligning with prompt/guardrail evaluation in LLM security.
-
DGF-Bench: A Benchmark for Simulating and Auditing Deception Against Multi-Agent Governance Boards
No Free Efficiency probes training efficiency vs vulnerability trade-offs across models; cross-domain relevance to security.
-
Self-Evolving Defense: Continual Security Policy Learning for LLM Agents
Item 1–RareTrap framework targets behavior estimation in black-box LLMs, tying into safety evaluation and policy considerations.
-
Do Defenses Against LLM Extraction Work Across Attacks? A Lifecycle Benchmark of Black-Box Model Extraction
Item 9 provides a lifecycle benchmark for model extraction defenses, directly addressing defenses and evaluative benchmarking.