Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
4,981 to 5,040 of 18,830
Deep Optics for Single-shot High-dynamic-range Imaging
Christopher A. Metzler, Hayato Ikoma, Yifan Peng +1
eess.IVcs.CVarXiv:1908.00620v12019CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models
Hao He, Ceyuan Yang, Shanchuan Lin +7
cs.CVarXiv:2503.10592v12025CoLT-Drive: Counterfactual Long-Tail Benchmarking and Knowledge-Preserving Adaptation for Driving Affordance Prediction
Zhengxu Tang, Guofeng Cui, Ziyu Gong +8
cs.CVcs.AIcs.CLarXiv:2609.00242v12026UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Jianke Zhang, Yanjiang Guo, Yucheng Hu +3
cs.CVcs.AIarXiv:2501.18867v32025Boundary Proposal Network for Two-Stage Natural Language Video Localization
Shaoning Xiao, Long Chen, Songyang Zhang +4
cs.CVarXiv:2103.08109v22021Unsupervised Adversarial Depth Estimation using Cycled Generative Networks
Andrea Pilzer, Dan Xu, Mihai Marian Puscas +2
cs.CVarXiv:1807.10915v12018RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning
Hao Gao, Shaoyu Chen, Bo Jiang +11
cs.CVcs.ROarXiv:2502.13144v22025Can Video World Models Track Unobserved World States?
Joonghyuk Shin, Yicong Hong, Jaesik Park +1
cs.CVarXiv:2608.30692v12026Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding
Xiang Fang, Daizong Liu, Wanlong Fang +4
cs.CVarXiv:2605.30742v12026Phrase Localization and Visual Relationship Detection with Comprehensive Image-Language Cues
Bryan A. Plummer, Arun Mallya, Christopher M. Cervantes +2
cs.CVarXiv:1611.06641v42016Multi-view PointNet for 3D Scene Understanding
Maximilian Jaritz, Jiayuan Gu, Hao Su
cs.CVarXiv:1909.13603v12019NeuMesh: Learning Disentangled Neural Mesh-based Implicit Field for Geometry and Texture Editing
Bangbang Yang, Chong Bao, Junyi Zeng +4
cs.CVcs.GRarXiv:2207.11911v12022Degradation-Aware Feature Perturbation for All-in-One Image Restoration
Xiangpeng Tian, Xiangyu Liao, Xiao Liu +2
cs.CVcs.AIarXiv:2505.12630v12025AlphaRAD: Grounded Zero-Shot Classification in Chest Radiology via $α$-Corrected Binary Cross Entropy and Factorized Latent Supervision
Jianzhong You, Yuan Gao, Chris McIntosh
cs.CVarXiv:2609.01757v12026Text2Earth: Unlocking Text-driven Remote Sensing Image Generation with a Global-Scale Dataset and a Foundation Model
Chenyang Liu, Keyan Chen, Rui Zhao +2
cs.CVarXiv:2501.00895v22025An Explainable 3D Residual Self-Attention Deep Neural Network FOR Joint Atrophy Localization and Alzheimer's Disease Diagnosis using Structural MRI
Xin Zhang, Liangxiu Han, Wenyong Zhu +2
eess.IVcs.CVarXiv:2008.04024v22020Framelet Representation of Tensor Nuclear Norm for Third-Order Tensor Completion
Tai-Xiang Jiang, Michael K. Ng, Xi-Le Zhao +1
eess.IVcs.CVarXiv:1909.06982v22019Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
Shulin Tian, Ruiqi Wang, Hongming Guo +7
cs.CVcs.AIarXiv:2506.13654v12025Ten Architectures, One Error: Shared Failure Modes in Hyperspectral Classification under Spatially Disjoint Evaluation
Ehsan Faghih, Fatemeh Ashrafi, Marguerite Moore +1
cs.CVcs.LGarXiv:2609.01786v12026Similarity-Guided Layer-Adaptive Vision Transformer for UAV Tracking
Chaocan Xue, Bineng Zhong, Qihua Liang +4
cs.CVarXiv:2503.06625v12025Learning Human Identity from Motion Patterns
Natalia Neverova, Christian Wolf, Griffin Lacey +4
cs.LGcs.CVcs.NEarXiv:1511.03908v42015Attention-GAN for Object Transfiguration in Wild Images
Xinyuan Chen, Chang Xu, Xiaokang Yang +1
cs.CVarXiv:1803.06798v12018STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution
Rui Xie, Yinhong Liu, Penghao Zhou +7
cs.CVarXiv:2501.02976v12025Cross-Modal Progressive Comprehension for Referring Segmentation
Si Liu, Tianrui Hui, Shaofei Huang +3
cs.CVcs.MMarXiv:2105.07175v12021TAPIP3D: Tracking Any Point in Persistent 3D Geometry
Bowei Zhang, Lei Ke, Adam W. Harley +1
cs.CVcs.LGarXiv:2504.14717v32025TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Michael S. Ryoo, AJ Piergiovanni, Anurag Arnab +2
cs.CVcs.LGarXiv:2106.11297v42021Optical Flow with Semantic Segmentation and Localized Layers
Laura Sevilla-Lara, Deqing Sun, Varun Jampani +1
cs.CVarXiv:1603.03911v22016VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception
Ziang Yan, Xinhao Li, Yinan He +6
cs.CVarXiv:2509.21100v12025Prior-Guided Implicit Neural Representations for Single-Subject Diffusion MRI Super-Resolution
Abdulkader Ghandoura, Marsil Zakour, William Consagra +1
eess.IVcs.CVarXiv:2609.00981v12026Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
Mohsen Gholami, Ahmad Rezaei, Zhou Weimin +4
cs.CVarXiv:2509.06266v22025ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents
Qiuchen Wang, Ruixue Ding, Zehui Chen +4
cs.CVcs.AIcs.CLarXiv:2502.18017v22025Channel-Wise Attention-Based Network for Self-Supervised Monocular Depth Estimation
Jiaxing Yan, Hong Zhao, Penghui Bu +1
cs.CVcs.AIcs.LGarXiv:2112.13047v12021BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
Yuming Li, Yikai Wang, Yuying Zhu +4
cs.CVcs.AIcs.LGarXiv:2509.06040v52025Weakly Supervised Deep Nuclei Segmentation Using Partial Points Annotation in Histopathology Images
Hui Qu, Pengxiang Wu, Qiaoying Huang +7
cs.CVarXiv:2007.05448v12020The Garden of Forking Paths: Towards Multi-Future Trajectory Prediction
Junwei Liang, Lu Jiang, Kevin Murphy +2
cs.CVarXiv:1912.06445v32019Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion
Shiyuan Yang, Liang Hou, Haibin Huang +5
cs.CVarXiv:2402.03162v22024Robust Object Modeling for Visual Tracking
Yidong Cai, Jie Liu, Jie Tang +1
cs.CVarXiv:2308.05140v12023D3D: Distilled 3D Networks for Video Action Recognition
Jonathan C. Stroud, David A. Ross, Chen Sun +2
cs.CVarXiv:1812.08249v22018Sparc3D: Sparse Representation and Construction for High-Resolution 3D Shapes Modeling
Zhihao Li, Yufei Wang, Heliang Zheng +2
cs.CVarXiv:2505.14521v32025NFormer: Robust Person Re-identification with Neighbor Transformer
Haochen Wang, Jiayi Shen, Yongtuo Liu +2
cs.CVcs.AIarXiv:2204.09331v12022YOLO Evolution: A Comprehensive Benchmark and Architectural Review of YOLOv12, YOLO11, and Their Previous Versions
Nidhal Jegham, Chan Young Koh, Marwan Abdelatti +1
cs.CVarXiv:2411.00201v42024Señorita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists
Bojia Zi, Penghui Ruan, Marco Chen +7
cs.CVarXiv:2502.06734v32025Improving the Diffusability of Autoencoders
Ivan Skorokhodov, Sharath Girish, Benran Hu +5
cs.CVcs.AIcs.LGarXiv:2502.14831v32025DuDoRNet: Learning a Dual-Domain Recurrent Network for Fast MRI Reconstruction with Deep T1 Prior
Bo Zhou, S. Kevin Zhou
eess.IVcs.CVarXiv:2001.03799v22020Looking Beyond the Scale: Do Surgical Skill Models Learn Transferable Representations Across Assessment Rubrics?
Hanna Hoffmann, Felix von Bechtolsheim, Stefanie Speidel +1
cs.CVcs.LGarXiv:2608.17519v12026ASSERT: Adaptive Stochastic Sampling for Robust Diffusion Models on Analog Compute-in-Memory Hardware
Yuannuo Feng, Yizhe Chen, Wenshuai Yao +4
cs.CVarXiv:2609.00955v12026Streaming4D: Accelerate 4D World Models via Block-wise Video Generation and Incremental Reconstruction
Xiaoyan Liu, Jiaxin Liu, Kangrui Li +1
cs.CVarXiv:2609.00610v12026Efficient Visual Pretraining with Contrastive Detection
Olivier J. Hénaff, Skanda Koppula, Jean-Baptiste Alayrac +3
cs.CVarXiv:2103.10957v22021Patchwork++: Fast and Robust Ground Segmentation Solving Partial Under-Segmentation Using 3D Point Cloud
Seungjae Lee, Hyungtae Lim, Hyun Myung
cs.ROcs.CVarXiv:2207.11919v22022Grounded Text-to-Image Synthesis with Attention Refocusing
Quynh Phung, Songwei Ge, Jia-Bin Huang
cs.CVarXiv:2306.05427v22023Vehicle Re-identification with Viewpoint-aware Metric Learning
Ruihang Chu, Yifan Sun, Yadong Li +3
cs.CVarXiv:1910.04104v12019PhysGen3D: Crafting a Miniature Interactive World from a Single Image
Boyuan Chen, Hanxiao Jiang, Shaowei Liu +4
cs.CVarXiv:2503.20746v12025GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation
Tong Wu, Guandao Yang, Zhibing Li +5
cs.CVarXiv:2401.04092v22024Dyn-3D: Unveiling and Resolving Ego-Motion Ambiguity in Vision-Language Models
Jiayu Ding, Zhuodong Liu, Lei Zhang +6
cs.CVarXiv:2609.01059v12026VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
Siyu Xu, Yunke Wang, Chenghao Xia +3
cs.ROcs.CVcs.LGarXiv:2502.02175v22025Two Stream LSTM: A Deep Fusion Framework for Human Action Recognition
Harshala Gammulle, Simon Denman, Sridha Sridharan +1
cs.CVarXiv:1704.01194v12017DMCP: Differentiable Markov Channel Pruning for Neural Networks
Shaopeng Guo, Yujie Wang, Quanquan Li +1
cs.CVcs.LGarXiv:2005.03354v22020Scaling Video Analytics on Constrained Edge Nodes
Christopher Canel, Thomas Kim, Giulio Zhou +5
cs.CVcs.LGcs.PFarXiv:1905.13536v12019DDT: Decoupled Diffusion Transformer
Shuai Wang, Zhi Tian, Weilin Huang +1
cs.CVcs.AIarXiv:2504.05741v22025MotionNet: Joint Perception and Motion Prediction for Autonomous Driving Based on Bird's Eye View Maps
Pengxiang Wu, Siheng Chen, Dimitris Metaxas
cs.CVarXiv:2003.06754v12020