惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

大猫的无限游戏
大猫的无限游戏
Webroot Blog
Webroot Blog
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
T
Threat Research - Cisco Blogs
V2EX - 技术
V2EX - 技术
L
LINUX DO - 热门话题
Google DeepMind News
Google DeepMind News
Recorded Future
Recorded Future
S
Schneier on Security
I
InfoQ
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
The GitHub Blog
The GitHub Blog
S
Security @ Cisco Blogs
O
OpenAI News
W
WeLiveSecurity
Vercel News
Vercel News
阮一峰的网络日志
阮一峰的网络日志
Simon Willison's Weblog
Simon Willison's Weblog
人人都是产品经理
人人都是产品经理
Cloudbric
Cloudbric
The Last Watchdog
The Last Watchdog
The Hacker News
The Hacker News
Google Online Security Blog
Google Online Security Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
GbyAI
GbyAI
NISL@THU
NISL@THU
T
Tailwind CSS Blog
V
Visual Studio Blog
PCI Perspectives
PCI Perspectives
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
Jina AI
Jina AI
D
DataBreaches.Net
B
Blog RSS Feed
N
News and Events Feed by Topic
N
News and Events Feed by Topic
H
Heimdal Security Blog
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
腾讯CDC
Latest news
Latest news
V
Vulnerabilities – Threatpost
Hacker News: Ask HN
Hacker News: Ask HN
WordPress大学
WordPress大学
V
V2EX
aimingoo的专栏
aimingoo的专栏
博客园 - 司徒正美
Apple Machine Learning Research
Apple Machine Learning Research
D
Darknet – Hacking Tools, Hacker News & Cyber Security
The Register - Security
The Register - Security
Help Net Security
Help Net Security

cs updates on arXiv.org

Beyond Binary Edits Robust Multimodal Knowledge Editing with Adversarial Subspace Alignment Agentic Proving for Program Verification MemAudit: Post-hoc Auditing of Poisoned Agent Memory via Causal Attribution and Structural Anomaly Detection OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents One Policy, Infinite NPCs: Persona-Traceable Shared RL Policies for Scalable Game Agents How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework Benchmarking Google Embeddings 2 against Open-Source Models for Multilingual Dense Retrieval and RAG Systems Structure-Guided Entity Resolution: Fine-Tuning LLMs for Robust Name Matching in Complex Linguistic Contexts Solving the Aircraft Disassembly Scheduling Problem Co-ReAct: Rubrics as Step-Level Collaborators for ReAct Agents CP or DP? Why Not Both: A Case Study in the Partial Shop Scheduling Problem Asking For An Old Friend: Diagnosing and Mitigating Temporal Failure Modes in LLM-based Statutory Question Answering EDGE-OPD: Internalizing Privileged Context with Evidence Guided On-Policy Distillation ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning SSDAU: Structured Semantic Data Augmentation for Joint Entity and Relation Extraction Naturalistic measure of social norms alignment Articulatory strategy as a source of variation in acoustic vowel dynamics When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems EquiSumm : A Gender Bias-Aware Framework for Inclusive Tweet Summarization Metacognition as Reward: Reinforcing LLM Reasoning via Knowledge and Regulation Signals From Correctness to Preference: A Framework for Personalized Agentic Reinforcement Learning Cultural Adaptation in Large Language Models for Political Discourse Emotion Recognition in Sign Language Conversation ClimateChat-300K: A Multi-Modal Facebook Dataset for Understanding Diverse Perspectives in Climate Communication AraHopeCorpus: Annotation Guidelines and Dataset for Hope Speech in Arabic Social Media Crisis Discourse Human-in-the-Loop Multi-Agent Ventilator Decision Support with Contextual Bandit Preference Learning Convergence Without Understanding: When Language Models Agree on Representations but Disagree on Reasoning DART: Semantic Recoverability for Structured Tool Agents Ontological Knowledge Blocks: Executable Compliance and Profile-Based Validation for Trustworthy AI Systems Parallel Context Compaction for Long-Horizon LLM Agent Serving When Is Next-Token Prediction Useful? Marginalization, Ergodicity, Mixture Identifiability, Local Sufficiency, RAG, Tools, and Programming Design and Report Benchmarks for Knowledge Work GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models Foundation Protocol: A Coordination Layer for Agentic Society AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery Hidden Human-Like Nature of Machine-Generated Texts: Theory and Detection Enhancement Self-Improving In-Context Learning Redrawing the AI Map: A Theory of Accountability Boundaries in Agentic Ecosystems Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks Fast-dDrive: Efficient Block-Diffusion VLM for Autonomous Driving Same Model, Different Weakness: How Language and Modality Reshape the Jailbreak Attack Surface in Frontier MLLMs When Symptoms Are Not Enough: Evidence-Weighting Patterns in Large Language Model Psychiatric Screening As X, Do Y: How Persona and Task Combine in Instruction-Tuned LLMs CoReVAD: A Contextual Reasoning Framework for Training-Free Video Anomaly Detection Inconsistency-aware Multimodal Schrödinger Bridge for Deepfake Localization Inductive Deductive Synthesis: Enabling AI to Generate Formally Verified Systems A Fine-Tuned BERT Classifier for Personal-Letter Titles in Late-Ming and Early-Qing Collected Works A Comparative Evaluation of Structural Topic Models and BERTopic for Short, Open-Ended Survey Responses PathCal: State-Aware Reflection-Marker Calibration for Efficient Reasoning The Efficiency Frontier: A Unified Framework for Cost-Performance Optimization in LLM Context Management Flow Mismatching: Unsupervised Anomaly Detection via Velocity Discrepancies in Flow Matching Models DFKI-MLT at SemEval-2026 TASK 7: Steering Multilingual Models Towards Cultural Knowledge RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering What Training Data Teaches RL Memory Agents: An Empirical Study of Curriculum Effects in Memory-Augmented QA Dithering Defense: Adversarial Robustness of Vision Foundation Models via Multi-Level Floyd-Steinberg Dithering Millimeter-wave Imaging for Anthropometric Body Measurement Model Collapse as Cultural Evolution DreamerNLplus: Interpretable Modeling of Mental Health Dynamics from Social Media Timelines using Hybrid Rule-Based and RAG Methods The TIME Machine: On The Power of Motion for Efficient Perception HawkesLLM: Semantic Uncertainty Propagation in Agentic Text Simulation Do Language Models Know What Not to Say? Causal Evidence for Statistical Preemption in LLMs Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection Sparse Autoencoders Map Brain-LLM Alignment onto Cortical Semantic Topography Brain-LLM Alignment Tracks Training Data, Not Typology The Deterministic Horizon: Impossibility Results as Design Specifications for Trustworthy AI Systems Scene Reconstruction as Mapping Priors for 3D Detection CoMoGen: COntrollable MOtion Dynamics and Interactions with Mask-Guided Video GENeration A Proactive Multi-Agent Dialogue Framework for Assessing Social Language Disorder Traits in Autism Memorization Dynamics of Fill-in-the-Middle Pretraining A Reproducible Universal Dependencies-Style Pipeline for Katharevousa Greek Parliamentary Text When AI Takes Sides on Questions of Faith: Persistent Asymmetries in AI-Mediated Faith Guidance Can AI Guess What You Know? Performance Comparison of Large Language Models for Human Domain Knowledge Estimation From Communication Logs Graph Alignment Topology as an Inductive Bias for Grounding Detection GazeBehavior Annotation Toolkit (GBAT): AI-powered toolkit for automatic annotation of egocentric eye-tracking and video data of child-caregiver interaction Improved Vision-to-Chart Buoy Association with Learned World-to-Image Projection Learnability-Informed Fine-Tuning of Diffusion Language Models RAS: Reflection-Augmented Scaling with In-Context Learning for Executable Cypher Query Generation VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding EVE-Agent: Evidence-Verifiable Self-Evolving Agents Suicide Risk Assessment from AI-powered Video Surveillance: An Interpretable Framework for Prevention in Metro Stations Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Mediative Fuzzy Logic: From Type-1 Foundations to Type-2, Type-3 and Quantum Extensions ImProver 2: Iteratively Self-Improving LMs for Neurosymbolic Proof Optimization Energy per Successful Goal: Goal-Level Energy Accounting for Agentic AI Systems GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation How Far Will They Go? Red-Teaming Online Influence with Large Language Models SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research RMA: an Agentic System for Research-Level Mathematical Problems NeuroNL2LTL: A Neurosymbolic Framework for Natural Language Translation of Linear Temporal Logic BOHM: Zero-Cost Hierarchical Attribution for Compound AI Systems GAGPO: Generalized Advantage Grouped Policy Optimization Knowledge Distillation for Low-Resource Open-source Text-to-SQL Model Query-Adaptive Semantic Chunking for Retrieval-Augmented Generation: A Dynamic Strategy with Contextual Window Expansion A Survey of Text and Speech Resources for Hausa and Fongbe: Availability, Quality, and Gaps for NLP Development Evaluating Large Language Models in a Complex Hidden Role Game AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering Memory-SAM: Human-Prompt-Free Tongue Segmentation via Retrieval-to-Prompt DIVER: Reinforced Diffusion Breaks Imitation Bottlenecks in End-to-End Autonomous Driving Transformer-Empowered Actor-Critic Reinforcement Learning for Sequence-Aware Service Function Chain Partitioning VerteNet -- A Multi-Context Hybrid CNN Transformer for Accurate Vertebral Landmark Localization in Lateral Spine DXA Images
PowerAgentBench-SS: A Benchmark for Agentic AI in Power System Steady-State Studies
[Submitted on 17 Jun 2026] · 2026-06-18 · via cs updates on arXiv.org

View PDF HTML (experimental)

Abstract:Power system benchmarks usually evaluate numerical solvers, prediction models, or sequential controllers. These benchmarks are necessary, but they do not directly test whether a Large Language Model (LLM) agent can execute an engineering workflow: inspect a grid case, select tools, call simulators, screen contingencies, propose admissible mitigations, validate results, and produce an auditable evidence trail. This paper introduces PowerAgentBench-SS, a steady-state benchmark framework for evaluating tool-using agents in power system operation and planning studies. The benchmark exposes public case data, action constraints, a tool API, and a validation budget to an agent, while a hidden evaluator recomputes physical validity and scores the submitted report. We define the agent interface, tool contract, evidence log, and risk-sensitive metrics, including submitted recall, evidence-backed recall, found recall, false-safe penalties, severity regret, residual violation score, action cost, tool-use efficiency, and workflow diagnostics. To make the framework concrete, we instantiate the protocol in a reproducible DC thermal N-2 contingency-search pilot on deterministic IEEE 39-bus operating-point variants, with scripted baselines, an LLM JSON-command adapter, three locally hosted Ollama LLM agents, and one OpenAI API agent. The results show why solver-only or answer-only evaluation is insufficient: agents are distinguished not only by top-contingency discovery, but also by validation-budget use, explicit submission, type coercions, duplicate validations, evidence-backed reporting, and mitigation behavior.

Submission history

From: Costas Mylonas [view email]
[v1] Wed, 17 Jun 2026 08:04:12 UTC (2,719 KB)