Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
5,041 to 5,100 of 18,821
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
Erik Daxberger, Nina Wenzel, David Griffiths +8
cs.CVcs.CLcs.LGarXiv:2503.13111v22025Radial Attention: $O(n\log n)$ Sparse Attention with Energy Decay for Long Video Generation
Xingyang Li, Muyang Li, Tianle Cai +11
cs.CVcs.AIcs.LGarXiv:2506.19852v22025DANNet: A One-Stage Domain Adaptation Network for Unsupervised Nighttime Semantic Segmentation
Xinyi Wu, Zhenyao Wu, Hao Guo +2
cs.CVarXiv:2104.10834v12021GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
Yuhan Wang, Siwei Yang, Bingchen Zhao +4
cs.CVarXiv:2507.21033v12025OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
Shihao Wang, Zhiding Yu, Xiaohui Jiang +6
cs.CVarXiv:2405.01533v22024WildGaussians: 3D Gaussian Splatting in the Wild
Jonas Kulhanek, Songyou Peng, Zuzana Kukelova +2
cs.CVarXiv:2407.08447v22024HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters
Yi Chen, Sen Liang, Zixiang Zhou +6
cs.CVarXiv:2505.20156v22025VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos
Xubin Ren, Lingrui Xu, Long Xia +3
cs.IRcs.AIcs.CVarXiv:2502.01549v12025Triangle Splatting for Real-Time Radiance Field Rendering
Jan Held, Renaud Vandeghen, Adrien Deliege +7
cs.CVarXiv:2505.19175v12025MINE: Towards Continuous Depth MPI with NeRF for Novel View Synthesis
Jiaxin Li, Zijian Feng, Qi She +3
cs.CVcs.GRcs.LGarXiv:2103.14910v32021PromptAD: Learning Prompts with only Normal Samples for Few-Shot Anomaly Detection
Xiaofan Li, Zhizhong Zhang, Xin Tan +4
cs.CVarXiv:2404.05231v22024CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification
Wei Li, Renshan Zhang, Rui Shao +2
cs.CVcs.ROarXiv:2508.21046v32025Multiple Sound Sources Localization from Coarse to Fine
Rui Qian, Di Hu, Heinrich Dinkel +3
cs.CVarXiv:2007.06355v22020Vision-Language-Guided Pseudo-Labels for Unsupervised Domain Adaptation in Semantic Segmentation for Waste Sorting
Udo Schlegel, Shubhangi, Gabriel Dax +3
cs.CVcs.AIcs.LGarXiv:2609.00898v12026DiT4SR: Taming Diffusion Transformer for Real-World Image Super-Resolution
Zheng-Peng Duan, Jiawei Zhang, Xin Jin +6
cs.CVarXiv:2503.23580v22025FlowDPS: Flow-Driven Posterior Sampling for Inverse Problems
Jeongsol Kim, Bryan Sangwoo Kim, Jong Chul Ye
cs.CVcs.AIcs.LGarXiv:2503.08136v12025Deep Face Super-Resolution with Iterative Collaboration between Attentive Recovery and Landmark Estimation
Cheng Ma, Zhenyu Jiang, Yongming Rao +2
cs.CVarXiv:2003.13063v12020Learning to Generate Images of Outdoor Scenes from Attributes and Semantic Layouts
Levent Karacan, Zeynep Akata, Aykut Erdem +1
cs.CVarXiv:1612.00215v12016ORSIm Detector: A Novel Object Detection Framework in Optical Remote Sensing Imagery Using Spatial-Frequency Channel Features
Xin Wu, Danfeng Hong, Jiaojiao Tian +3
cs.CVarXiv:1901.07925v22019Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation
Sucheng Ren, Qihang Yu, Ju He +3
cs.CVarXiv:2502.20388v22025Multi-View Image Generation from a Single-View
Bo Zhao, Xiao Wu, Zhi-Qi Cheng +3
cs.CVcs.MMarXiv:1704.04886v42017Multimodal Alignment and Fusion: A Survey
Songtao Li, Hao Tang
cs.CVarXiv:2411.17040v22024LiDAR-Camera Calibration using 3D-3D Point correspondences
Ankit Dhall, Kunal Chelani, Vishnu Radhakrishnan +1
cs.ROcs.CVarXiv:1705.09785v12017End-to-End Race Driving with Deep Reinforcement Learning
Maximilian Jaritz, Raoul de Charette, Marin Toromanoff +2
cs.CVcs.ROarXiv:1807.02371v22018A Review on Explainability in Multimodal Deep Neural Nets
Gargi Joshi, Rahee Walambe, Ketan Kotecha
cs.AIcs.CVarXiv:2105.07878v22021ChatCAD: Interactive Computer-Aided Diagnosis on Medical Image using Large Language Models
Sheng Wang, Zihao Zhao, Xi Ouyang +2
cs.CVeess.IVarXiv:2302.07257v12023VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
Yichao Shen, Fangyun Wei, Zhiying Du +5
cs.ROcs.AIcs.CVarXiv:2512.06963v12025Improving Clinical Target Volume Segmentation Accuracy using Anatomical Priors and Active Learning for the AGITG TOPGEAR Clinical Trial
Phillip Chlap, Mark Lee, Trevor Leong +11
physics.med-phcs.CVarXiv:2609.03186v12026Deep Learning in Automated Power Line Inspection: A Review
Md. Ahasan Atick Faisal, Imene Mecheter, Yazan Qiblawey +3
cs.CVeess.IVarXiv:2502.07826v12025Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers
Wei Pang, Kevin Qinghong Lin, Xiangru Jian +2
cs.CVcs.AIcs.CLarXiv:2505.21497v22025ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
Wenxuan Song, Ziyang Zhou, Han Zhao +7
cs.ROcs.CVarXiv:2508.10333v12025LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models
Shenghao Fu, Qize Yang, Qijie Mo +5
cs.CVarXiv:2501.18954v12025Building Extraction at Scale using Convolutional Neural Network: Mapping of the United States
Hsiuhan Lexie Yang, Jiangye Yuan, Dalton Lunga +3
cs.CVarXiv:1805.08946v12018Insert Anything: Image Insertion via In-Context Editing in DiT
Wensong Song, Hong Jiang, Zongxing Yang +2
cs.CVarXiv:2504.15009v12025Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
Size Wu, Wenwei Zhang, Lumin Xu +6
cs.CVarXiv:2503.21979v22025CAMEL: A Weakly Supervised Learning Framework for Histopathology Image Segmentation
Gang Xu, Zhigang Song, Zhuo Sun +6
eess.IVcs.CVcs.LGarXiv:1908.10555v12019Quad-networks: unsupervised learning to rank for interest point detection
Nikolay Savinov, Akihito Seki, Lubor Ladicky +2
cs.CVcs.LGcs.NEarXiv:1611.07571v22016Rethinking Weakly-supervised Video Temporal Grounding From a Game Perspective
Xiang Fang, Zeyu Xiong, Wanlong Fang +7
cs.CVcs.AIarXiv:2605.26441v12026HoliTom: Holistic Token Merging for Fast Video Large Language Models
Kele Shao, Keda Tao, Can Qin +3
cs.CVarXiv:2505.21334v32025PixNerd: Pixel Neural Field Diffusion
Shuai Wang, Ziteng Gao, Chenhui Zhu +2
cs.CVarXiv:2507.23268v22025O-CNN: Octree-based Convolutional Neural Networks for 3D Shape Analysis
Peng-Shuai Wang, Yang Liu, Yu-Xiao Guo +2
cs.CVarXiv:1712.01537v12017Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
Haoyu Wu, Diankun Wu, Tianyu He +4
cs.CVcs.AIarXiv:2507.07982v22025Spatially Aware World Action Model via Geometric Latent Diffusion
Javier Alejandro Lopetegui Gonzalez, Paul Pacaud, Cordelia Schmid
cs.CVcs.ROarXiv:2609.02531v12026MST: Masked Self-Supervised Transformer for Visual Representation
Zhaowen Li, Zhiyang Chen, Fan Yang +8
cs.CVarXiv:2106.05656v220213D Face Morphable Models "In-the-Wild"
James Booth, Epameinondas Antonakos, Stylianos Ploumpis +3
cs.CVarXiv:1701.05360v12017Image-Grounded Conversations: Multimodal Context for Natural Question and Response Generation
Nasrin Mostafazadeh, Chris Brockett, Bill Dolan +4
cs.CLcs.AIcs.CVarXiv:1701.08251v22017BRISC: Annotated Dataset for Brain Tumor Segmentation and Classification
Amirreza Fateh, Yasin Rezvani, Sara Moayedi +4
eess.IVcs.CVarXiv:2506.14318v52025Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing
Yusu Qian, Eli Bocek-Rivele, Liangchen Song +5
cs.CVcs.CLcs.LGarXiv:2510.19808v12025moco: Fast Motion Correction for Calcium Imaging
Alexander Dubbs, James Guevara, Darcy S. Peterka +1
cs.CVarXiv:1506.06039v12015An unscented Kalman filter method for real time input-parameter-state estimation
Marios Impraimakis, Andrew W. Smyth
eess.SPcs.AIcs.CVarXiv:2511.02717v12025Unrolled Optimization with Deep Priors
Steven Diamond, Vincent Sitzmann, Felix Heide +1
cs.CVarXiv:1705.08041v22017Large-scale and Fine-grained Vision-language Pre-training for Enhanced CT Image Understanding
Zhongyi Shui, Jianpeng Zhang, Weiwei Cao +8
cs.CVarXiv:2501.14548v12025FPGA/DNN Co-Design: An Efficient Design Methodology for IoT Intelligence on the Edge
Cong Hao, Xiaofan Zhang, Yuhong Li +5
cs.CVarXiv:1904.04421v12019Depth-Based 3D Hand Pose Estimation: From Current Achievements to Future Goals
Shanxin Yuan, Guillermo Garcia-Hernando, Bjorn Stenger +21
cs.CVarXiv:1712.03917v22017Characterizing Text Branch Sensitivity in Medical Vision-Language Segmentation via Evidence Decoupling
Ziquan Liu, Zhewei Zhu, Xuyang Shi
cs.CVarXiv:2609.02663v12026DriveAdapter: Breaking the Coupling Barrier of Perception and Planning in End-to-End Autonomous Driving
Xiaosong Jia, Yulu Gao, Li Chen +3
cs.ROcs.CVarXiv:2308.00398v22023LaST-SR: Laplace-Inspired Steady-Transient Complex-Frequency Decomposition for Single Image Super-Resolution
Linhao Li, Zhaojie Pan, Langkun Chen
cs.CVarXiv:2609.02063v12026Uni-Sign: Toward Unified Sign Language Understanding at Scale
Zecheng Li, Wengang Zhou, Weichao Zhao +3
cs.CVarXiv:2501.15187v32025Towards Causal VQA: Revealing and Reducing Spurious Correlations by Invariant and Covariant Semantic Editing
Vedika Agarwal, Rakshith Shetty, Mario Fritz
cs.CVcs.CLcs.LGarXiv:1912.07538v32019Automatic Detection of Knee Joints and Quantification of Knee Osteoarthritis Severity using Convolutional Neural Networks
Joseph Antony, Kevin McGuinness, Kieran Moran +1
cs.CVarXiv:1703.09856v12017