Browse all

arXiv research papers — all

All arXiv papers (AI/ML focus) — filter and search by category and year.

Clear

174 results

PaperCategoryAuthorsDate
Rubrics on Trial: Evolving Rubrics from a Single Query via Synth... ↗ cs.CL 8 Jul 16, 2026
OmniaBench: Benchmarking General AI Agents Across Diverse Scenar... ↗ cs.CL 16 Jul 16, 2026
Latent Trajectory Discrimination for AI-Generated Text Detection ↗ cs.CL 8 Jul 16, 2026
Show Me How You Reason and I'll Tell You Who You Are: Reasoning... ↗ cs.CL 4 Jul 16, 2026
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforce... ↗ cs.CL 11 Jul 16, 2026
Dialogue Summarization with Emotion Dynamics Using Topic- and Pa... ↗ cs.CL 3 Jul 16, 2026
CoTu at EXACT 2026: Neuro-Symbolic Reasoning for Transparent Edu... ↗ cs.CL 6 Jul 16, 2026
Testing proactive agents in live Docker containers — UniClawBenc... Explained cs.CL 7 Jul 9, 2026
Validity of LLMs as data annotators: AMALIA on authority ↗ cs.CL 1 Jul 9, 2026
You do not need a frontier model to verify citations — calibrati... Explained cs.CL 6 Jul 9, 2026
WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide... ↗ cs.CL 11 Jul 9, 2026
UltraX: Refining Pre-Training Data at Scale with Adaptive Progra... ↗ cs.CL 12 Jul 9, 2026
DominoTree: Conditional Tree-Structured Drafting with Domino for... ↗ cs.CL 2 Jul 9, 2026
It Takes a MAESTRO To Prune Bad Experts ↗ cs.CL 3 Jul 9, 2026
When the Judge Changes, So Does the Measurement: Auditing LLM-as... ↗ cs.CL 3 Jul 9, 2026
Cross-seed explainability using Procrustes-conditioned Joint End... ↗ cs.CL 2 Jul 9, 2026
Two Axes of LLM Abstention: Answer Correctness and Question Answ... ↗ cs.CL 1 Jul 9, 2026
Detecting Ladder Logic Bombs in IEC 61131-3 PLC Programs using E... ↗ cs.CL 3 Jul 9, 2026
When Synthetic Speech Is All You Have: Better Call GRPO ↗ cs.CL 8 Jul 9, 2026
Prompt Compression via Activation Aggregation ↗ cs.CL 4 Jul 9, 2026
Large-Language-Models-as-a-Judge in Theory-Agnostic Adaptive Met... ↗ cs.CL 9 Jul 9, 2026
Echoes Across Vietnam's Highlands, Delta, and Coast: A Multiling... ↗ cs.CL 5 Jul 9, 2026
Grounded Event Extraction from SEC 8-K Filings with a Fine-Grain... ↗ cs.CL 5 Jul 9, 2026
TypeProbe: Recovering Type Representations from Hidden States of... ↗ cs.CL 2 Jul 9, 2026
XALPHA: A Memory-Driven AI Quant Researcher for Hypothesis-to-Co... ↗ cs.CL 4 Jul 9, 2026
Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment ↗ cs.CL 2 Jul 9, 2026
LACUNA: A Testbed for Evaluating Localization Precision for LLM... ↗ cs.CL 5 Jul 2, 2026
Reasoning LLM Improves Speaker Recognition in Long-form TV Drama... ↗ cs.CL 9 Jul 2, 2026
Visually Grounded Self-Reflection for Vision-Language Models via... ↗ cs.CL 3 Jul 2, 2026
Audio-Based Understanding of Audiobook Narration Appeal ↗ cs.CL 3 Jul 2, 2026
Will Scaling Improve Social Simulation with LLMs? ↗ cs.CL 6 Jul 2, 2026
Language Models as Measurement Apparatus for Culture ↗ cs.CL 1 Jul 2, 2026
The Future of NLP may not be at NLP Conferences: Scholarly Migra... ↗ cs.CL 1 Jul 2, 2026
Know Your Source: A Public Knowledge Store for Media Background... ↗ cs.CL 3 Jul 2, 2026
HULAT2 at MER-TRANS 2026: Governed Multi-Agent Simplification fo... ↗ cs.CL 4 Jul 2, 2026
World Wide Models: Literary Tools for Cultural AI ↗ cs.CL 1 Jul 2, 2026
On the Role of Directionality in Structural Generalization ↗ cs.CL 1 Jul 2, 2026
CheckRLM: Effective Knowledge-Thought Coherence Checking in Retr... ↗ cs.CL 11 Jul 2, 2026
BamiBERT: A New BERT-based Language Model for Vietnamese ↗ cs.CL 4 Jul 2, 2026
Challenges and Recommendations for LLMs-as-a-Judge in Multilingu... ↗ cs.CL 6 Jul 2, 2026
Unlocking Speech-Text Compositional Powers: Instruction-Followin... ↗ cs.CL 4 Jul 2, 2026
Mapping Political-Elite Networks in Europe with a Multilingual J... ↗ cs.CL 2 Jun 25, 2026
Empowering GUI Agents via Autonomous Experience Exploration and... ↗ cs.CL 6 Jun 25, 2026
LLM-Based Examination of Eligibility Criteria from Securities Pr... ↗ cs.CL 3 Jun 25, 2026
Beyond Surface Forms: A Comprehensive, Mechanism-Oriented Taxono... ↗ cs.CL 4 Jun 25, 2026
Multilingual Reasoning Cascades Need More Context ↗ cs.CL 5 Jun 25, 2026
How Surprising Is Historical Italian to Language Models? Tokeniz... ↗ cs.CL 1 Jun 25, 2026
LMs as Task-Specific Knowledge Bases: An Interpretability Analys... ↗ cs.CL 3 Jun 25, 2026
Bridging Talk and Thought: Understanding Dialogue Dynamics Acros... ↗ cs.CL 4 Jun 25, 2026
CARVE: Content-Aware Recurrent with Value Efficiency for Chunk-P... ↗ cs.CL 1 Jun 25, 2026

Source: official U.S. government open data. This is an organized index, not an official U.S. government site. "Explained" links to our summary page; otherwise links go to the official primary source.

Disclaimer: This site independently summarizes and classifies information based on official data sources. Always verify the latest and accurate information with the official sources. Content on finance, health, legal, and security is information, not advice. This site is not an official website of the U.S. government.