Browse all

arXiv research papers — all

All arXiv papers (AI/ML focus) — filter and search by category and year.

Clear

334 results

PaperCategoryAuthorsDate
ReToken: One Token to Improve Vision-Language Models for Visual... ↗ cs.CV 6 Jul 30, 2026
ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engin... ↗ cs.CV 16 Jul 30, 2026
PhiZero: A World Model Built Around Physical Language ↗ cs.CV 7 Jul 30, 2026
Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusio... ↗ cs.CV 12 Jul 30, 2026
Beacon: Knowing When and How to Perform Agentic Visual Reasoning ↗ cs.CV 14 Jul 30, 2026
VAD: Attributing Visual Evidence for Target Reconstruction in Mu... ↗ cs.CV 12 Jul 30, 2026
MixFrag: Fragility-Guided Mixed-Precision Post-Training Quantiza... ↗ cs.CV 3 Jul 30, 2026
ROAD: Reciprocal-Objective Alignment of Discriminative Semantics... ↗ cs.CV 8 Jul 30, 2026
Finding Change in Satellite Archives from Text: How to Combine B... ↗ cs.CV 3 Jul 30, 2026
MIND: Multimodal Intent-Driven Network via Diffusion Transformer... ↗ cs.CV 6 Jul 30, 2026
ScaFE: Data-Efficient Scar Classification with LLM-Generated Cli... ↗ cs.CV 2 Jul 30, 2026
MarkushGlyph and OCSRGlyph: Improved Chemical Structure Recognit... ↗ cs.CV 4 Jul 30, 2026
What to Remove, What to Preserve: Dual-Ambiguity Rectification f... ↗ cs.CV 9 Jul 30, 2026
Beyond Frame Selection: Generative Latent Evidence Aggregation f... ↗ cs.CV 6 Jul 30, 2026
RefCaptioner: Multi-Reference Image-Grounded Video Captioning ↗ cs.CV 19 Jul 30, 2026
AuricularWorld: Hierarchical Action-Guided World Modeling for Fi... ↗ cs.CV 9 Jul 30, 2026
Towards Real-Time PixOOD: Efficient Anomaly Segmentation for Aut... ↗ cs.CV 4 Jul 30, 2026
Can Vision-Language Models Reason about AI Edits in Images? ↗ cs.CV 4 Jul 30, 2026
VisualRouter: Query-Grounded Visual Sampling for Long Video Unde... ↗ cs.CV 8 Jul 30, 2026
ViewMind3D: Modular View-Aware Inference for Training-Free 3D-QA ↗ cs.CV 6 Jul 30, 2026
Negative controls reveal volume-driven confounding in radiomics... ↗ cs.CV 9 Jul 30, 2026
Large scale cross-regional remote sensing flood monitoring frame... ↗ cs.CV 9 Jul 30, 2026
Hand-Object Interaction in the Age of Large Foundation Models:Re... ↗ cs.CV 9 Jul 30, 2026
Explaining Image Similarity with Automatically Extracted Concept... ↗ cs.CV 4 Jul 30, 2026
ShadowDancer: Teaching Video World Models Any Action by Learning... ↗ cs.CV 3 Jul 30, 2026
Capturing Token Tendencies for Training-Free Token Pruning in Mu... ↗ cs.CV 7 Jul 30, 2026
Same Branches, Different Trees: A Bifurcation Connectedness Metr... ↗ cs.CV 6 Jul 30, 2026
AdaAnchor4D: Anchor-Conditioned Spatiotemporal Feature Aggregati... ↗ cs.CV 9 Jul 30, 2026
ObjectStream: Latent Objects as Memory Anchors for Streaming Vid... ↗ cs.CV 11 Jul 30, 2026
MonoVoc: Decoupling Geometry and Semantics for Lightweight Monoc... ↗ cs.CV 4 Jul 30, 2026
Filling the Pareto-Optimal Front for Affordance Segmentation on... ↗ cs.CV 5 Jul 30, 2026
Beyond Visual Ambiguity: Guiding Robust Monocular Depth Estimati... ↗ cs.CV 5 Jul 30, 2026
MSCM-net: A hyperspectral image classiffcation method based on m... ↗ cs.CV 7 Jul 30, 2026
Theia: Large-Scale Multimodal Captioning and Automated Validatio... ↗ cs.CV 4 Jul 30, 2026
TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting ↗ cs.CV 8 Jul 30, 2026
Space2Ground 2.0: A Multi-Source Dataset and Framework for Agric... ↗ cs.CV 5 Jul 30, 2026
EgoGenesis: Egocentric World-Action Modeling with Online Anchore... ↗ cs.CV 12 Jul 30, 2026
FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Ima... ↗ cs.CV 5 Jul 30, 2026
Scaling Vision-Language Models Is Not Enough to Mitigate Bias ↗ cs.CV 3 Jul 30, 2026
Think with Extra-Image: A Farmland Segmentation Agent Driven by... ↗ cs.CV 7 Jul 30, 2026
3D-Aware VLMs with Implicit and Explicit Geometries ↗ cs.CV 7 Jul 23, 2026
Streaming Multi-Agent Autoregressive Diffusion Model with World... ↗ cs.CV 5 Jul 23, 2026
Unified Video Dense Prediction from Disjoint Data ↗ cs.CV 5 Jul 23, 2026
Inference-Time Scaling of Diffusion Models via Progressive Seed... ↗ cs.CV 2 Jul 23, 2026
GraphVid: Interactive Graph-Controllable Video Generation ↗ cs.CV 8 Jul 23, 2026
Synthetic data generation framework for quality control automati... ↗ cs.CV 4 Jul 23, 2026
Self-Supervised Learning of Structured Dynamics from Videos ↗ cs.CV 3 Jul 23, 2026
Scene Parameter Saliency via Differentiable Light Transport ↗ cs.CV 2 Jul 23, 2026
Visual Contrastive Self-Distillation ↗ cs.CV 7 Jul 23, 2026
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals... ↗ cs.CV 14 Jul 23, 2026

Source: official U.S. government open data. This is an organized index, not an official U.S. government site. "Explained" links to our summary page; otherwise links go to the official primary source.

Disclaimer: This site independently summarizes and classifies information based on official data sources. Always verify the latest and accurate information with the official sources. Content on finance, health, legal, and security is information, not advice. This site is not an official website of the U.S. government.