Browse all

arXiv research papers — all

All arXiv papers (AI/ML focus) — filter and search by category and year.

Clear

334 results

PaperCategoryAuthorsDate
Frequency-Structured Field Learning for Light-Field Disparity Es... ↗ cs.CV 3 Jul 16, 2026
VideoChat3: Fully Open Video MLLM for Efficient and Generalist V... ↗ cs.CV 27 Jul 16, 2026
Still image and spatial-temporal tomato data enabling detection,... ↗ cs.CV 10 Jul 16, 2026
Benchmarking Face Recognition without Real Faces ↗ cs.CV 5 Jul 16, 2026
TanGO: Training-Free 3D Editing via Tangent-Space Guidance and O... ↗ cs.CV 5 Jul 16, 2026
FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with T... ↗ cs.CV 2 Jul 16, 2026
Selectivity Drives Efficiency: Dataset Pruning for Visual Place... ↗ cs.CV 6 Jul 16, 2026
Rotational Motion-Induced Error Compensation for Phase-Shifting... ↗ cs.CV 4 Jul 16, 2026
Physics-Informed Diffusion for Biomechanically Plausible 3D Sign... ↗ cs.CV 5 Jul 16, 2026
Blurring Modal Boundaries: A Unified Survey from Single- to Mult... ↗ cs.CV 6 Jul 16, 2026
An LLM-Based Automatic Sportscast Solution for Robot Soccer Matc... ↗ cs.CV 6 Jul 16, 2026
TAMF-VTON: Texture-Aware Mask-Free Virtual Try-On via High-Fidel... ↗ cs.CV 5 Jul 16, 2026
Rare Concept Generation via Counterfactual Inference in Diffusio... ↗ cs.CV 4 Jul 16, 2026
Clean-Reference Streaming Detection of Lens Occlusion and Photom... ↗ cs.CV 3 Jul 16, 2026
On the Disagreement in Perturbation-based xAI -- Benchmarking Pe... ↗ cs.CV 2 Jul 16, 2026
FoMoVLA: Bridging Visual Foresight and Motion Guidance for Visio... ↗ cs.CV 12 Jul 16, 2026
GeoDetect: Geometric Adversarial Detection for VLPs ↗ cs.CV 5 Jul 16, 2026
Zero-shot monocular depth in 6.1M parameters — "ZipDepth," light... Explained cs.CV 4 Jul 9, 2026
Wat3R: Underwater 3D Geometry Learning without Annotations ↗ cs.CV 7 Jul 9, 2026
LongE2V: Long-Horizon Event-based Video Reconstruction, Predicti... ↗ cs.CV 7 Jul 9, 2026
Geometry and Gradient-based Partitioning for Panoramic Outdoor R... ↗ cs.CV 10 Jul 9, 2026
OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step A... ↗ cs.CV 9 Jul 9, 2026
Enhancing In-context Panoramic Generation via Geometric-aware Pr... ↗ cs.CV 5 Jul 9, 2026
OpenCoF: Learning to Reason Through Video Generation ↗ cs.CV 5 Jul 9, 2026
WaspMOT: A Benchmark for Long-Term Multi-Object Tracking of Tric... ↗ cs.CV 7 Jul 9, 2026
Pose-to-Biomechanics: Bridging 3D Human Pose Estimation and Biom... ↗ cs.CV 2 Jul 9, 2026
LTM: Large-scale Terrain Model for Wildfire-prone Landscapes ↗ cs.CV 5 Jul 9, 2026
HumanForge: A Human-Centric Deepfake Video Benchmark with Multi-... ↗ cs.CV 5 Jul 9, 2026
SAM-MT: Real-Time Interactive Multi-Target Video Segmentation ↗ cs.CV 3 Jul 9, 2026
Multi-Resolution Feature Stem for Diabetic Retinopathy lesion se... ↗ cs.CV 2 Jul 9, 2026
Do Transformations Reveal the Truth? Generative Residual Learnin... ↗ cs.CV 5 Jul 9, 2026
When Structured Sparse Autoencoders Learn Consistent Concepts Ac... ↗ cs.CV 3 Jul 9, 2026
Switch-Reasoner: Learn When to Think in Multitask Mixtures via R... ↗ cs.CV 10 Jul 9, 2026
VocaDet: Sample-Driven Open-Vocabulary Object Detection and Segm... ↗ cs.CV 1 Jul 9, 2026
Whareformer: Learning to Track What is Where in Long Egocentric... ↗ cs.CV 5 Jul 9, 2026
Beyond wheelchairs and blindfolds: Investigating disability ster... ↗ cs.CV 3 Jul 9, 2026
Do Egocentric Video-Language Models Capture Both Hand- and Objec... ↗ cs.CV 5 Jul 9, 2026
CT-CLIP Representations for Multimodal Lung Cancer Survival Pred... ↗ cs.CV 7 Jul 9, 2026
Cognitive-structured Multimodal Agent for Multimodal Understandi... ↗ cs.CV 6 Jul 9, 2026
VEGAS: Human-Aligned Video Caption Evaluation via Gaze ↗ cs.CV 8 Jul 9, 2026
Predicting Viticulture Potential through an Ensemble of U-Net an... ↗ cs.CV 3 Jul 9, 2026
DeltaV: Thinking with Visual State Updates in Unified Large Mult... ↗ cs.CV 9 Jul 9, 2026
Track2Map: Online Deformable SLAM with Motion-Aware Pose Optimiz... ↗ cs.CV 9 Jul 9, 2026
Swapping Faces, Saving Features: A Dual-Purpose Pipeline for Ped... ↗ cs.CV 2 Jul 9, 2026
Attribute Retrieving for Open-Vocabulary Endoscopic Compositiona... ↗ cs.CV 6 Jul 9, 2026
Classical versus Deep Mirror-Symmetry Scoring: A Benchmark of Th... ↗ cs.CV 1 Jul 9, 2026
WCog-VLA: A Dual-Level World-Cognitive Vision-Language-Action Mo... ↗ cs.CV 8 Jul 9, 2026
Texture Representations in Deep Vision Models: Comparing CNNs, V... ↗ cs.CV 4 Jul 9, 2026
ARGUS: Accelerated, Robust, General, and Unsupervised Cell Track... ↗ cs.CV 4 Jul 9, 2026
Enhancing the KidSat Model: Integrating Geographical Encoding an... ↗ cs.CV 7 Jul 9, 2026

Source: official U.S. government open data. This is an organized index, not an official U.S. government site. "Explained" links to our summary page; otherwise links go to the official primary source.

Disclaimer: This site independently summarizes and classifies information based on official data sources. Always verify the latest and accurate information with the official sources. Content on finance, health, legal, and security is information, not advice. This site is not an official website of the U.S. government.