Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

5,041 to 5,100 of 18,821

  1. MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs

    Erik Daxberger, Nina Wenzel, David Griffiths +8

    cs.CVcs.CLcs.LGarXiv:2503.13111v22025
  2. Radial Attention: $O(n\log n)$ Sparse Attention with Energy Decay for Long Video Generation

    Xingyang Li, Muyang Li, Tianle Cai +11

    cs.CVcs.AIcs.LGarXiv:2506.19852v22025
  3. DANNet: A One-Stage Domain Adaptation Network for Unsupervised Nighttime Semantic Segmentation

    Xinyi Wu, Zhenyao Wu, Hao Guo +2

    cs.CVarXiv:2104.10834v12021
  4. GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset

    Yuhan Wang, Siwei Yang, Bingchen Zhao +4

    cs.CVarXiv:2507.21033v12025
  5. OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning

    Shihao Wang, Zhiding Yu, Xiaohui Jiang +6

    cs.CVarXiv:2405.01533v22024
  6. WildGaussians: 3D Gaussian Splatting in the Wild

    Jonas Kulhanek, Songyou Peng, Zuzana Kukelova +2

    cs.CVarXiv:2407.08447v22024
  7. HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

    Yi Chen, Sen Liang, Zixiang Zhou +6

    cs.CVarXiv:2505.20156v22025
  8. VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos

    Xubin Ren, Lingrui Xu, Long Xia +3

    cs.IRcs.AIcs.CVarXiv:2502.01549v12025
  9. Triangle Splatting for Real-Time Radiance Field Rendering

    Jan Held, Renaud Vandeghen, Adrien Deliege +7

    cs.CVarXiv:2505.19175v12025
  10. MINE: Towards Continuous Depth MPI with NeRF for Novel View Synthesis

    Jiaxin Li, Zijian Feng, Qi She +3

    cs.CVcs.GRcs.LGarXiv:2103.14910v32021
  11. PromptAD: Learning Prompts with only Normal Samples for Few-Shot Anomaly Detection

    Xiaofan Li, Zhizhong Zhang, Xin Tan +4

    cs.CVarXiv:2404.05231v22024
  12. CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification

    Wei Li, Renshan Zhang, Rui Shao +2

    cs.CVcs.ROarXiv:2508.21046v32025
  13. Multiple Sound Sources Localization from Coarse to Fine

    Rui Qian, Di Hu, Heinrich Dinkel +3

    cs.CVarXiv:2007.06355v22020
  14. Vision-Language-Guided Pseudo-Labels for Unsupervised Domain Adaptation in Semantic Segmentation for Waste Sorting

    Udo Schlegel, Shubhangi, Gabriel Dax +3

    cs.CVcs.AIcs.LGarXiv:2609.00898v12026
  15. DiT4SR: Taming Diffusion Transformer for Real-World Image Super-Resolution

    Zheng-Peng Duan, Jiawei Zhang, Xin Jin +6

    cs.CVarXiv:2503.23580v22025
  16. FlowDPS: Flow-Driven Posterior Sampling for Inverse Problems

    Jeongsol Kim, Bryan Sangwoo Kim, Jong Chul Ye

    cs.CVcs.AIcs.LGarXiv:2503.08136v12025
  17. Deep Face Super-Resolution with Iterative Collaboration between Attentive Recovery and Landmark Estimation

    Cheng Ma, Zhenyu Jiang, Yongming Rao +2

    cs.CVarXiv:2003.13063v12020
  18. Learning to Generate Images of Outdoor Scenes from Attributes and Semantic Layouts

    Levent Karacan, Zeynep Akata, Aykut Erdem +1

    cs.CVarXiv:1612.00215v12016
  19. ORSIm Detector: A Novel Object Detection Framework in Optical Remote Sensing Imagery Using Spatial-Frequency Channel Features

    Xin Wu, Danfeng Hong, Jiaojiao Tian +3

    cs.CVarXiv:1901.07925v22019
  20. Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation

    Sucheng Ren, Qihang Yu, Ju He +3

    cs.CVarXiv:2502.20388v22025
  21. Multi-View Image Generation from a Single-View

    Bo Zhao, Xiao Wu, Zhi-Qi Cheng +3

    cs.CVcs.MMarXiv:1704.04886v42017
  22. Multimodal Alignment and Fusion: A Survey

    Songtao Li, Hao Tang

    cs.CVarXiv:2411.17040v22024
  23. LiDAR-Camera Calibration using 3D-3D Point correspondences

    Ankit Dhall, Kunal Chelani, Vishnu Radhakrishnan +1

    cs.ROcs.CVarXiv:1705.09785v12017
  24. End-to-End Race Driving with Deep Reinforcement Learning

    Maximilian Jaritz, Raoul de Charette, Marin Toromanoff +2

    cs.CVcs.ROarXiv:1807.02371v22018
  25. A Review on Explainability in Multimodal Deep Neural Nets

    Gargi Joshi, Rahee Walambe, Ketan Kotecha

    cs.AIcs.CVarXiv:2105.07878v22021
  26. ChatCAD: Interactive Computer-Aided Diagnosis on Medical Image using Large Language Models

    Sheng Wang, Zihao Zhao, Xi Ouyang +2

    cs.CVeess.IVarXiv:2302.07257v12023
  27. VideoVLA: Video Generators Can Be Generalizable Robot Manipulators

    Yichao Shen, Fangyun Wei, Zhiying Du +5

    cs.ROcs.AIcs.CVarXiv:2512.06963v12025
  28. Improving Clinical Target Volume Segmentation Accuracy using Anatomical Priors and Active Learning for the AGITG TOPGEAR Clinical Trial

    Phillip Chlap, Mark Lee, Trevor Leong +11

    physics.med-phcs.CVarXiv:2609.03186v12026
  29. Deep Learning in Automated Power Line Inspection: A Review

    Md. Ahasan Atick Faisal, Imene Mecheter, Yazan Qiblawey +3

    cs.CVeess.IVarXiv:2502.07826v12025
  30. Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers

    Wei Pang, Kevin Qinghong Lin, Xiangru Jian +2

    cs.CVcs.AIcs.CLarXiv:2505.21497v22025
  31. ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver

    Wenxuan Song, Ziyang Zhou, Han Zhao +7

    cs.ROcs.CVarXiv:2508.10333v12025
  32. LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

    Shenghao Fu, Qize Yang, Qijie Mo +5

    cs.CVarXiv:2501.18954v12025
  33. Building Extraction at Scale using Convolutional Neural Network: Mapping of the United States

    Hsiuhan Lexie Yang, Jiangye Yuan, Dalton Lunga +3

    cs.CVarXiv:1805.08946v12018
  34. Insert Anything: Image Insertion via In-Context Editing in DiT

    Wensong Song, Hong Jiang, Zongxing Yang +2

    cs.CVarXiv:2504.15009v12025
  35. Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

    Size Wu, Wenwei Zhang, Lumin Xu +6

    cs.CVarXiv:2503.21979v22025
  36. CAMEL: A Weakly Supervised Learning Framework for Histopathology Image Segmentation

    Gang Xu, Zhigang Song, Zhuo Sun +6

    eess.IVcs.CVcs.LGarXiv:1908.10555v12019
  37. Quad-networks: unsupervised learning to rank for interest point detection

    Nikolay Savinov, Akihito Seki, Lubor Ladicky +2

    cs.CVcs.LGcs.NEarXiv:1611.07571v22016
  38. Rethinking Weakly-supervised Video Temporal Grounding From a Game Perspective

    Xiang Fang, Zeyu Xiong, Wanlong Fang +7

    cs.CVcs.AIarXiv:2605.26441v12026
  39. HoliTom: Holistic Token Merging for Fast Video Large Language Models

    Kele Shao, Keda Tao, Can Qin +3

    cs.CVarXiv:2505.21334v32025
  40. PixNerd: Pixel Neural Field Diffusion

    Shuai Wang, Ziteng Gao, Chenhui Zhu +2

    cs.CVarXiv:2507.23268v22025
  41. O-CNN: Octree-based Convolutional Neural Networks for 3D Shape Analysis

    Peng-Shuai Wang, Yang Liu, Yu-Xiao Guo +2

    cs.CVarXiv:1712.01537v12017
  42. Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling

    Haoyu Wu, Diankun Wu, Tianyu He +4

    cs.CVcs.AIarXiv:2507.07982v22025
  43. Spatially Aware World Action Model via Geometric Latent Diffusion

    Javier Alejandro Lopetegui Gonzalez, Paul Pacaud, Cordelia Schmid

    cs.CVcs.ROarXiv:2609.02531v12026
  44. MST: Masked Self-Supervised Transformer for Visual Representation

    Zhaowen Li, Zhiyang Chen, Fan Yang +8

    cs.CVarXiv:2106.05656v22021
  45. 3D Face Morphable Models "In-the-Wild"

    James Booth, Epameinondas Antonakos, Stylianos Ploumpis +3

    cs.CVarXiv:1701.05360v12017
  46. Image-Grounded Conversations: Multimodal Context for Natural Question and Response Generation

    Nasrin Mostafazadeh, Chris Brockett, Bill Dolan +4

    cs.CLcs.AIcs.CVarXiv:1701.08251v22017
  47. BRISC: Annotated Dataset for Brain Tumor Segmentation and Classification

    Amirreza Fateh, Yasin Rezvani, Sara Moayedi +4

    eess.IVcs.CVarXiv:2506.14318v52025
  48. Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing

    Yusu Qian, Eli Bocek-Rivele, Liangchen Song +5

    cs.CVcs.CLcs.LGarXiv:2510.19808v12025
  49. moco: Fast Motion Correction for Calcium Imaging

    Alexander Dubbs, James Guevara, Darcy S. Peterka +1

    cs.CVarXiv:1506.06039v12015
  50. An unscented Kalman filter method for real time input-parameter-state estimation

    Marios Impraimakis, Andrew W. Smyth

    eess.SPcs.AIcs.CVarXiv:2511.02717v12025
  51. Unrolled Optimization with Deep Priors

    Steven Diamond, Vincent Sitzmann, Felix Heide +1

    cs.CVarXiv:1705.08041v22017
  52. Large-scale and Fine-grained Vision-language Pre-training for Enhanced CT Image Understanding

    Zhongyi Shui, Jianpeng Zhang, Weiwei Cao +8

    cs.CVarXiv:2501.14548v12025
  53. FPGA/DNN Co-Design: An Efficient Design Methodology for IoT Intelligence on the Edge

    Cong Hao, Xiaofan Zhang, Yuhong Li +5

    cs.CVarXiv:1904.04421v12019
  54. Depth-Based 3D Hand Pose Estimation: From Current Achievements to Future Goals

    Shanxin Yuan, Guillermo Garcia-Hernando, Bjorn Stenger +21

    cs.CVarXiv:1712.03917v22017
  55. Characterizing Text Branch Sensitivity in Medical Vision-Language Segmentation via Evidence Decoupling

    Ziquan Liu, Zhewei Zhu, Xuyang Shi

    cs.CVarXiv:2609.02663v12026
  56. DriveAdapter: Breaking the Coupling Barrier of Perception and Planning in End-to-End Autonomous Driving

    Xiaosong Jia, Yulu Gao, Li Chen +3

    cs.ROcs.CVarXiv:2308.00398v22023
  57. LaST-SR: Laplace-Inspired Steady-Transient Complex-Frequency Decomposition for Single Image Super-Resolution

    Linhao Li, Zhaojie Pan, Langkun Chen

    cs.CVarXiv:2609.02063v12026
  58. Uni-Sign: Toward Unified Sign Language Understanding at Scale

    Zecheng Li, Wengang Zhou, Weichao Zhao +3

    cs.CVarXiv:2501.15187v32025
  59. Towards Causal VQA: Revealing and Reducing Spurious Correlations by Invariant and Covariant Semantic Editing

    Vedika Agarwal, Rakshith Shetty, Mario Fritz

    cs.CVcs.CLcs.LGarXiv:1912.07538v32019
  60. Automatic Detection of Knee Joints and Quantification of Knee Osteoarthritis Severity using Convolutional Neural Networks

    Joseph Antony, Kevin McGuinness, Kieran Moran +1

    cs.CVarXiv:1703.09856v12017