论文
每日采集自 arXiv 与 Hugging Face Papers,未成稿的也全部保留。
- Correct Is Not Governed: Provenance Integrity in Agentic Workflows
- MindMemOS: A Portable and Self-Evolving Memory Operating Layer for AI Agents
- Mimicry without understanding: the origins of decision bias in large language models
- Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies
- ReflectFact: Self-Reflective Agents for Improving Comprehension and Reasoning in Multi-Hop Fact Verification
- Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents
- Thought-Aware KV Cache Compaction for Reasoning via Adaptive Attention Matching
- Learning to Adapt Cross-Domain Preferences via Meta-LoRA for LLM Personalization
- Designing AI Pipelines for Decision-Ready ITSM Intelligence
- $\varepsilon$-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution
- New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMs
- Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents
- LLMs Are Not Good Strategists, Yet Memory-Enhanced Agency Boosts Reasoning
- Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting
- General Probabilities of Causation with Causal Knowledge
- CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives
- Position: The Alignment Community is Unintentionally Building a Censor's Toolkit
- Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese
- Trie Automata for Constrained Decoding over Large Finite Sets
- SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries
- Lines and Ladders: A Context-Aware Multi-Agent Framework for Large-Scale Retail Price Taxonomy
- The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis
- PROVE-RT: Generating Mechanized Theorem Prover Scripts for Real-Time Systems using LLMs
- ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs
- Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories
- Query Timing Produces Opposite Positional Biases Between LLMs and Humans
- AI and Consumer Rights in India Working Paper
- Are you Talking Logic to Me? Assessing Language Models Syllogistic Reasoning Capabilities
- Novels generated by language models show compressed formal variation
- DiG-bench: Discovery in Games
- SDAM: Structure-Difference-Aware Memory Evolution for Complex Text-to-SQL
- Reliability-Aware Sexism Detection: Combining DPO with Annotator Agreement and Token-Level Confidence Scoring
- LoRA-Diffusion: Parameter-Efficient Fine-Tuning via Low-Rank Trajectory Decomposition
- Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
- Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces
- Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists
- StorySpark: Module-wise Evolutionary Search for Story Premise Generation
- Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning
- Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction
- CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence
- Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues
- Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence
- On the Expressive Power of Transformers
- When Explanations Betray Backdoors: Black-Box Auditing for Language Model Classifiers
- Intensional Anaphora
- Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs
- PatientAct: Theory-Grounded Mental Health Client Simulation
- DIVE: Unlocking Self-Improvement in Frozen Language Models Through Diversity-Driven Skill Evolution
- Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models
- Predicting consumer-technology ownership without a diffusion history
- Large Language Models Pass the History Exam But Miss the <<History>>: A Polish High School Exit Exam Matura Benchmark
- CABS+: Efficient and Scalable Model Merging via Conflict-Aware Sparsification and Adaptive Weight Allocation
- From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options
- Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents
- The "Knowledge-Behavior Gap" in Cultural Taboo Safety of Large Language Models
- ReconSpan: Reconstruction-Guided Adaptive Latent Tokenization
- Predictive Memory Localization: Forecasting Selective Intervention Paths from Internal Signals
- LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning
- Class-Structure Preservation Beats Diversity: A Comprehensive Benchmark of Text Augmentation Methods for Imbalanced Text Classification
- @skills: Attention is all you have
- From Refuse to Richness: Rubric Rewards for Long-Form Hallucination Reinforcement Learning
- AQuA: Recursively Self-Improving Quantitative Trading Research Agents
- FastThaiG2P: Lightning-fast Thai Grapheme-to-phoneme Conversion for Voice Agent Pipelines
- Steering the Language Axis: From Linear Decodability to Causal Control
- On Measuring Semantic Preservation in Legal Ontology Learning
- HC-RAG: Evidence-Centric Retrieval-Augmented Generation over Heterogeneous Financial Filings
- What Drives LLM Self-Reflection? A Controlled Ablation of Uncertainty Routing in Armed Conflict Forecasting
- Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation
- AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement
- Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition
- Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance
- Can Spectral-Clipping Enable Better Learning While Forgetting Less for Low-Rank Adaptation?
- Vision-Language Models are Fragile Multilingual Associators
- Position: Reasoning is a Learnable Rule-Based Process
- ViTOED: A Dataset for Target-Oriented Emotion Detection on Vietnamese Social Media Texts
- Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments
- ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval
- Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing
- Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection
- Research Assistant: AstraZeneca's Agentic System for R&D
- Lost in Compaction: Evaluating Side-Constraint Loss under Context Compaction
- Better, Faster, Stronger: Programmatic Skill Learning Best Reduces Agent Cost
- Detecting a Route Flip Is Easier Than Knowing Whether to Fix It: Causal Route-Mediated Damage in Quantized Mixture-of-Experts
- A Forced-Structure Reduction and Verifiable Bounds for Conway's 99-Graph
- GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation
- Adaptive Hybrid Particle Swarm Optimization with Gradient Descent
- Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval
- When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use
- Conformity Mitigations in Large Language Models Lie on a Single Resistance-Receptivity Frontier
- Towards the Harness of Embodied Agents
- Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs
- Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop
- AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research
- Forecasting Side Effects of Activation Steering
- Measure, Don't Optimize: Forecasting Recovery in LLM Unlearning
- EvoGraph-Mem: Failure-Aware Editable Graph Memory for Long-Term Language Agents
- Towards Sustainable Learning in Online Education: A Reinforcement Learning Approach
- Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration
- The Off-Support Barrier: Why Semantic Safety Constraints Are Not Learning-Problem Invariants, and What Follows for Prior Design, Containment, and Verification
- Harnessing agent memory to build lifelong AI partners for materials scientists
- LazyTrain: Limited-resource Allocation toward Zero-waste Yield Optimization in Large Language Model Training
- Glance, Scrutinize, and Think: Advancing Video Anomaly Detection from Training-Free to Agentic Reasoning
- LLMs in Process Diagram Engineering: From Optimal PFDs to Validated P&IDs
- From Monolithic to Modular: Segment-level Automatic Prompt Optimization
- Ripple-Pivot Search: Active Parallel Decoding for Diffusion Large Language Models
- DonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech Recognition
- When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs
- Stigma and Support in Online Sexual Violence Narratives on Reddit
- Symbolic Machine Learning for Vapor-Liquid Equilibrium Prediction in Cx-N2 Binary Mixtures
- VQ-bench: A Composable Vector Quantization Framework
- Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems
- When the Knowledge Base Becomes the Gold Standard: Measuring Resource-Shared Evaluation Loops in Entity-Level Machine Translation
- Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations
- Language-Conditional Dequantization: Recovering What Quantization Steals from Non-English Languages
- Is Convergence Inevitable? Tracing Output Homogeneity Back to Base Models
- LabelFusion-TS: Fusing Large Language Models, Transformer Encoders, and Financial Time Series for Monetary-Policy Stance Classification
- Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability
- LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs
- InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk
- From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate
- Causal Structure is Inducible but Functionally Decoupled: The Routing/Readout Boundary of a Typed Mechanism Library
- Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction
- CT-$\Delta$Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models
- A Conceptual Framework for Refining Influence Knowledge from Simulation Evidence in Cyber-Physical Systems
- Self-Evolving Embodied Agents via Skill-Harness Evolution
- Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces
- BEST-KAG: Enhancing Question Answering of Building Engineering Standards with Multimodal Knowledge Graph Modeling and Large Language Model
- Principal Trait Analysis: Towards Deriving "Skills" in Human-AI Collaboration
- Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment
- Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets
- AgonAlpha: Autonomous Alpha Discovery via Prompt Economy and Scalable Agentic Search
- Synchronizing Beliefs with Second-Order Theory-of-Mind in Human-Autonomy Teams (Extended Version)
- Easper: An Accessible ASR Pipeline for Language Documentation
- Hybrid Gated Attention
- Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs
- MaSRead: Content-Addressed Reading of Replicated Latent Stores
- Locating and Controlling Implicit Personalization in Large Language Models
- ODE-Based Transformer Decoders for Iterative Sign Language Translation
- Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression
- Geometry-aware Incremental Neural Operator for Long-Horizon PDE prediction
- Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study
- TELLME: Test-Enhanced Learning for Language Model Enrichment
- Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence
- Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes
- Gloss-Free Representation Learning for Cross-Dataset Sign Spotting
- AWARe: Mitigating Catastrophic Forgetting via Activation-Weighted Adaptive REtention
- TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation
- Semantic Lenia: Emergence of Homeostatic Solitons within the Semantic Space of Large Language Models
- Cutting AI Datacenter Energy with Reinforcement Learning: Measured Power Control of LLM Training from One GPU to the Fleet
- The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance