Browse all

arXiv research papers — all

All arXiv papers (AI/ML focus) — filter and search by category and year.

Clear

334 results

PaperCategoryAuthorsDate
Don't Settle at the Mode! Mitigating Diversity Collapse in Pretr... ↗ cs.CV 4 Jun 25, 2026
PhysiFormer: Learning to Simulate Mechanics in World Space ↗ cs.CV 3 Jun 25, 2026
RayPE: Ray-Space Positional Encoding for 3D-Aware Video Generati... ↗ cs.CV 6 Jun 25, 2026
SAM2Matting: Generalized Image and Video Matting ↗ cs.CV 4 Jun 25, 2026
RoPEMover: Depth-Aware Object Relocation via Positional Embeddin... ↗ cs.CV 4 Jun 25, 2026
Not All Actions Are Equal: Rethinking Conditioning for Dexterous... ↗ cs.CV 10 Jun 25, 2026
OctoSense: Self-Supervised Learning for Multimodal Robot Percept... ↗ cs.CV 5 Jun 25, 2026
ViQ: Text-Aligned Visual Quantized Representations at Any Resolu... ↗ cs.CV 8 Jun 25, 2026
See & Sniff: Learning Visuo-Olfactory Representations ↗ cs.CV 5 Jun 25, 2026
Sculpting NeRF Geometry: Human-Preference Fine-Tuning of a 3D-Aw... ↗ cs.CV 3 Jun 25, 2026
Exact and Deterministic Patch Descriptor Retrieval via Hierarchi... ↗ cs.CV 1 Jun 25, 2026
CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Ches... ↗ cs.CV 6 Jun 25, 2026
SatSplatDiff: Geometry-preserving generative refinement for high... ↗ cs.CV 3 Jun 25, 2026
LISA: Likelihood Score Alignment for Visual-condition Controllab... ↗ cs.CV 7 Jun 25, 2026
HarmVideoBench: Benchmarking Harmful Video Understanding in Larg... ↗ cs.CV 16 Jun 25, 2026
Safe Autoregressive Image Generation with Iterative Self-Improvi... ↗ cs.CV 4 Jun 25, 2026
FlameVQA: A Physically-Grounded UAV Wildfire VQA Benchmark with... ↗ cs.CV 4 Jun 25, 2026
TMP: Tree-structured Mixed-policy Pruning for Large-scale Image... ↗ cs.CV 13 Jun 25, 2026
SubdivAR: Autoregressive Next-Scale Prediction for Neural Mesh S... ↗ cs.CV 10 Jun 25, 2026
Pseudo-Text-Conditioned 3D Grounding DINO for Organ Localization... ↗ cs.CV 6 Jun 25, 2026
PanoImager: Geometry-Guided Novel View Synthesis and Reconstruct... ↗ cs.CV 2 Jun 25, 2026
On-board Remote-Sensing Foundation Models for Unsupervised Chang... ↗ cs.CV 1 Jun 25, 2026
Event-Aware Instructed Assistant for Referring Video Segmentatio... ↗ cs.CV 4 Jun 25, 2026
Unison: Benchmarking Unified Multimodal Models via Synergistic U... ↗ cs.CV 4 Jun 25, 2026
Geometric Gradient Rectification for Safe Open-Set Semi-Supervis... ↗ cs.CV 7 Jun 25, 2026
Computer Vision for MOBA Analytics: A Dataset and Baseline for V... ↗ cs.CV 5 Jun 25, 2026
Scaling Multi-Reference Image Generation with Dynamic Reward Opt... ↗ cs.CV 9 Jun 25, 2026
TraMP-LLaMA: Generative Interpretability with Decoupled Instruct... ↗ cs.CV 5 Jun 25, 2026
Focusing on What Matters: Saliency-Harnessing Accurate Routing f... ↗ cs.CV 7 Jun 25, 2026
PortraitGen: Exemplar-Driven GRPO with Dual-Reward Guidance for... ↗ cs.CV 8 Jun 25, 2026
PhysRAG: Enhancing Physics-Awareness in Video Generation via Ret... ↗ cs.CV 5 Jun 25, 2026
Qwen-Image-Agent: Bridging the Context Gap in Real-World Image G... ↗ cs.CV 21 Jun 25, 2026
Confidence-Aware Tool Orchestration for Robust Video Understandi... ↗ cs.CV 3 Jun 25, 2026
Tractography-Driven Synthetic Data Generation for Fiber Bundle S... ↗ cs.CV 7 Jun 25, 2026
Modeling Local, Global, and Cross-Modal Context in Multimodal 3D... ↗ cs.CV 5 Jun 25, 2026
JanusMesh: Fast and Zero-Shot 3D Visual Illusion Generation via... ↗ cs.CV 4 Jun 18, 2026
TimeProVe: Propose, then Verify for Efficient Long Video Tempora... ↗ cs.CV 5 Jun 18, 2026
UNIEGO: Proxies as Mediators for Unified Egocentric Video Repres... ↗ cs.CV 5 Jun 18, 2026
Thinking in Boxes: 3D Editing in Real Images Made Easy ↗ cs.CV 7 Jun 18, 2026
Current World Models Lack a Persistent State Core ↗ cs.CV 11 Jun 18, 2026
SSD: Spatially Speculative Decoding Accelerates Autoregressive I... ↗ cs.CV 4 Jun 18, 2026
CalTennis: Large Multi-View Tennis Video Dataset and Benchmark o... ↗ cs.CV 5 Jun 18, 2026
The FID Lottery: Quantifying Hidden Randomness in Generative-Mod... ↗ cs.CV 3 Jun 18, 2026
VisDom: Sparse Novel View Synthesis with Visible Domain Constrai... ↗ cs.CV 6 Jun 18, 2026
SARLO-80: Worldwide Slant SAR Language Optic Dataset 80cm ↗ cs.CV 5 Jun 18, 2026
HumanScale: Egocentric Human Video Can Outperform Real-Robot Dat... ↗ cs.CV 22 Jun 18, 2026
S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intellig... ↗ cs.CV 13 Jun 18, 2026
FreeStyle: Free Control of Style-Content Dual-Reference Generati... ↗ cs.CV 13 Jun 18, 2026
How Fragile Are Training-Free AI-Generated Image Detectors? A Co... ↗ cs.CV 2 Jun 18, 2026
Scalable Training of Spatially Grounded 2D Vision-Language Model... ↗ cs.CV 7 Jun 18, 2026

Source: official U.S. government open data. This is an organized index, not an official U.S. government site. "Explained" links to our summary page; otherwise links go to the official primary source.

Disclaimer: This site independently summarizes and classifies information based on official data sources. Always verify the latest and accurate information with the official sources. Content on finance, health, legal, and security is information, not advice. This site is not an official website of the U.S. government.