Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
4,921 to 4,980 of 18,817
Humanoid Policy ~ Human Policy
Ri-Zhao Qiu, Shiqi Yang, Xuxin Cheng +12
cs.ROcs.AIcs.CVarXiv:2503.13441v32025CARRADA Dataset: Camera and Automotive Radar with Range-Angle-Doppler Annotations
A. Ouaknine, A. Newson, J. Rebut +2
cs.CVarXiv:2005.01456v62020Lepard: Learning partial point cloud matching in rigid and deformable scenes
Yang Li, Tatsuya Harada
cs.CVarXiv:2111.12591v22021Boundary Unlearning
Min Chen, Weizhuo Gao, Gaoyang Liu +2
cs.CVarXiv:2303.11570v12023Efficient Diffusion on Region Manifolds: Recovering Small Objects with Compact CNN Representations
Ahmet Iscen, Giorgos Tolias, Yannis Avrithis +2
cs.CVarXiv:1611.05113v32016Revisiting IM2GPS in the Deep Learning Era
Nam Vo, Nathan Jacobs, James Hays
cs.CVarXiv:1705.04838v12017TokenHSI: Unified Synthesis of Physical Human-Scene Interactions through Task Tokenization
Liang Pan, Zeshi Yang, Zhiyang Dou +5
cs.CVarXiv:2503.19901v22025Deep-Learned Collision Avoidance Policy for Distributed Multi-Agent Navigation
Pinxin Long, Wenxi Liu, Jia Pan
cs.AIcs.CVcs.ROarXiv:1609.06838v22016Collaborative Unsupervised Domain Adaptation for Medical Image Diagnosis
Yifan Zhang, Ying Wei, Qingyao Wu +4
cs.CVeess.IVarXiv:2007.07222v12020Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
Shanshan Zhao, Xinjie Zhang, Jintao Guo +9
cs.CVarXiv:2505.02567v62025VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning
Zhangyang Qi, Zhixiong Zhang, Yizhou Yu +2
cs.CVarXiv:2506.17221v22025SpatialTrust: A Benchmark for Environmental Risk Recognition in Secure Authentication
Junbin Lu, Hsiang-Wei Huang, Saesha Wadhwa +2
cs.CRcs.CVarXiv:2608.29489v12026Iterative Normalization: Beyond Standardization towards Efficient Whitening
Lei Huang, Yi Zhou, Fan Zhu +2
cs.CVcs.LGarXiv:1904.03441v120193DRS: MLLMs Need 3D-Aware Representation Supervision for Scene Understanding
Xiaohu Huang, Jingjing Wu, Qunyi Xie +1
cs.CVarXiv:2506.01946v220253D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer
Jiajun Deng, Tianyu He, Li Jiang +3
cs.CVarXiv:2501.01163v22025Beyond Blind Compliance: Benchmarking Task Verification in OCR Reasoning
Yue Zhou, Yuan Wu, Yi Chang
cs.CVarXiv:2609.00232v12026Video Object Segmentation with Adaptive Feature Bank and Uncertain-Region Refinement
Yongqing Liang, Xin Li, Navid Jafari +1
cs.CVarXiv:2010.07958v12020Learning Latent Action World Models In The Wild
Quentin Garrido, Tushar Nagarajan, Basile Terver +3
cs.AIcs.CVarXiv:2601.05230v22026Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
Haoji Zhang, Yiqin Wang, Yansong Tang +3
cs.CVarXiv:2506.23825v22025A Unified approach for Conventional Zero-shot, Generalized Zero-shot and Few-shot Learning
Shafin Rahman, Salman H. Khan, Fatih Porikli
cs.CVarXiv:1706.08653v22017SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration
Jianyi Wang, Zhijie Lin, Meng Wei +5
cs.CVarXiv:2501.01320v42025ARF: Artistic Radiance Fields
Kai Zhang, Nick Kolkin, Sai Bi +4
cs.CVarXiv:2206.06360v12022Asymmetric Loss Functions and Deep Densely Connected Networks for Highly Imbalanced Medical Image Segmentation: Application to Multiple Sclerosis Lesion Detection
Seyed Raein Hashemi, Seyed Sadegh Mohseni Salehi, Deniz Erdogmus +3
cs.CVarXiv:1803.11078v42018Neural Prompt Search
Yuanhan Zhang, Kaiyang Zhou, Ziwei Liu
cs.CVcs.AIcs.LGarXiv:2206.04673v22022Dr. Splat: Directly Referring 3D Gaussian Splatting via Direct Language Embedding Registration
Kim Jun-Seong, GeonU Kim, Kim Yu-Ji +3
cs.CVarXiv:2502.16652v12025MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
Wenyi Hong, Yean Cheng, Zhuoyi Yang +6
cs.CVarXiv:2501.02955v22025Actor and Observer: Joint Modeling of First and Third-Person Videos
Gunnar A. Sigurdsson, Abhinav Gupta, Cordelia Schmid +2
cs.CVarXiv:1804.09627v12018SpatialVID: A Large-Scale Video Dataset with Spatial Annotations
Jiahao Wang, Yufeng Yuan, Rujie Zheng +12
cs.CVarXiv:2509.09676v22025Diffusion-4K: Ultra-High-Resolution Image Synthesis with Latent Diffusion Models
Jinjin Zhang, Qiuyu Huang, Junjie Liu +2
cs.CVarXiv:2503.18352v22025ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
Zhongyi Zhou, Yichen Zhu, Junjie Wen +2
cs.ROcs.AIcs.CVarXiv:2505.21906v22025Generated Faces in the Wild: Quantitative Comparison of Stable Diffusion, Midjourney and DALL-E 2
Ali Borji
cs.CVarXiv:2210.00586v22022VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
Delong Chen, Mustafa Shukor, Theo Moutakanni +7
cs.CVarXiv:2512.10942v22025P3Depth: Monocular Depth Estimation with a Piecewise Planarity Prior
Vaishakh Patil, Christos Sakaridis, Alexander Liniger +1
cs.CVcs.AIcs.LGarXiv:2204.02091v12022Cultural Moment Benchmark: Evaluating Video Cultural Reasoning and Grounding in Southeast Asia
Burak Satar, Zhixin Ma, Cheng Yu-Tong +3
cs.CVcs.AIcs.CLarXiv:2608.23065v12026Robust Video Content Alignment and Compensation for Rain Removal in a CNN Framework
Jie Chen, Cheen-Hau Tan, Junhui Hou +2
cs.CVarXiv:1803.10433v12018Improving Vision-Language-Action Model with Online Reinforcement Learning
Yanjiang Guo, Jianke Zhang, Xiaoyu Chen +4
cs.ROcs.CVcs.LGarXiv:2501.16664v12025SparseFlex: High-Resolution and Arbitrary-Topology 3D Shape Modeling
Xianglong He, Zi-Xin Zou, Chia-Hao Chen +6
cs.CVarXiv:2503.21732v12025ReNFT: Repairing Mode Collapse in Reward Post-Training via Internal Probability-Mass Recalibration
Yuchen Bao, Chao Wen, Haowei Wang +10
cs.LGcs.AIcs.CVarXiv:2609.00061v12026PolaFormer: Polarity-aware Linear Attention for Vision Transformers
Weikang Meng, Yadan Luo, Xin Li +2
cs.CVcs.AIarXiv:2501.15061v22025OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
Shenghai Yuan, Xianyi He, Yufan Deng +5
cs.CVcs.AIarXiv:2505.20292v42025XNect: Real-time Multi-Person 3D Motion Capture with a Single RGB Camera
Dushyant Mehta, Oleksandr Sotnychenko, Franziska Mueller +7
cs.CVcs.GRarXiv:1907.00837v22019Real-Time Object Detection Meets DINOv3
Shihua Huang, Yongjie Hou, Longfei Liu +2
cs.CVarXiv:2509.20787v42025Attentive Relational Networks for Mapping Images to Scene Graphs
Mengshi Qi, Weijian Li, Zhengyuan Yang +2
cs.CVarXiv:1811.10696v22018NeuSG: Neural Implicit Surface Reconstruction with 3D Gaussian Splatting Guidance
Hanlin Chen, Chen Li, Yunsong Wang +1
cs.CVarXiv:2312.00846v22023GlobalBuildingAtlas: An Open Global and Complete Dataset of Building Polygons, Heights and LoD1 3D Models
Xiao Xiang Zhu, Sining Chen, Fahong Zhang +2
cs.CVarXiv:2506.04106v12025Convolution in Convolution for Network in Network
Yanwei Pang, Manli Sun, Xiaoheng Jiang +1
cs.CVarXiv:1603.06759v12016Beyond Textual Chain-of-Thought: A Survey on Action-Grounded Reasoning in Autonomous Driving
Zhengxu Tang, Xiaozhou Zhang, Guofeng Cui +10
cs.CVcs.CLcs.ROarXiv:2609.01659v12026Deep Optics for Single-shot High-dynamic-range Imaging
Christopher A. Metzler, Hayato Ikoma, Yifan Peng +1
eess.IVcs.CVarXiv:1908.00620v12019CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models
Hao He, Ceyuan Yang, Shanchuan Lin +7
cs.CVarXiv:2503.10592v12025CoLT-Drive: Counterfactual Long-Tail Benchmarking and Knowledge-Preserving Adaptation for Driving Affordance Prediction
Zhengxu Tang, Guofeng Cui, Ziyu Gong +8
cs.CVcs.AIcs.CLarXiv:2609.00242v12026UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Jianke Zhang, Yanjiang Guo, Yucheng Hu +3
cs.CVcs.AIarXiv:2501.18867v32025Boundary Proposal Network for Two-Stage Natural Language Video Localization
Shaoning Xiao, Long Chen, Songyang Zhang +4
cs.CVarXiv:2103.08109v22021Unsupervised Adversarial Depth Estimation using Cycled Generative Networks
Andrea Pilzer, Dan Xu, Mihai Marian Puscas +2
cs.CVarXiv:1807.10915v12018RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning
Hao Gao, Shaoyu Chen, Bo Jiang +11
cs.CVcs.ROarXiv:2502.13144v22025Can Video World Models Track Unobserved World States?
Joonghyuk Shin, Yicong Hong, Jaesik Park +1
cs.CVarXiv:2608.30692v12026Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding
Xiang Fang, Daizong Liu, Wanlong Fang +4
cs.CVarXiv:2605.30742v12026Phrase Localization and Visual Relationship Detection with Comprehensive Image-Language Cues
Bryan A. Plummer, Arun Mallya, Christopher M. Cervantes +2
cs.CVarXiv:1611.06641v42016Multi-view PointNet for 3D Scene Understanding
Maximilian Jaritz, Jiayuan Gu, Hao Su
cs.CVarXiv:1909.13603v12019NeuMesh: Learning Disentangled Neural Mesh-based Implicit Field for Geometry and Texture Editing
Bangbang Yang, Chong Bao, Junyi Zeng +4
cs.CVcs.GRarXiv:2207.11911v12022Degradation-Aware Feature Perturbation for All-in-One Image Restoration
Xiangpeng Tian, Xiangyu Liao, Xiao Liu +2
cs.CVcs.AIarXiv:2505.12630v12025