Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

6,061 to 6,120 of 18,916

  1. Osprey: Pixel Understanding with Visual Instruction Tuning

    Yuqian Yuan, Wentong Li, Jian Liu +5

    cs.CVarXiv:2312.10032v42023
  2. VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents

    Rui Meng, Ziyan Jiang, Ye Liu +10

    cs.CVcs.CLarXiv:2507.04590v12025
  3. MambaIRv2: Attentive State Space Restoration

    Hang Guo, Yong Guo, Yaohua Zha +5

    eess.IVcs.CVcs.LGarXiv:2411.15269v22024
  4. SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

    Xiyao Wang, Zhengyuan Yang, Chao Feng +6

    cs.CVarXiv:2504.07934v32025
  5. MonoPerfCap: Human Performance Capture from Monocular Video

    Weipeng Xu, Avishek Chatterjee, Michael Zollhöfer +4

    cs.CVcs.GRarXiv:1708.02136v22017
  6. A multicenter benchmark and clinically structured metric for coronary CTA report generation

    Zhiyu Ye, Yue Sun, Limiao Zou +7

    cs.CVarXiv:2609.00909v12026
  7. CMRVision: A Foundation Model for Cardiac MR Image Analysis

    Athira J. Jacob, Puneet Sharma, Daniel Rueckert

    cs.CVarXiv:2609.01308v12026
  8. Pose Estimation for Non-Cooperative Spacecraft Rendezvous Using Convolutional Neural Networks

    Sumant Sharma, Connor Beierle, Simone D'Amico

    cs.CVarXiv:1809.07238v12018
  9. Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks

    Tilman Räuker, Anson Ho, Stephen Casper +1

    cs.LGcs.AIcs.CLarXiv:2207.13243v62022
  10. TS2-Net: Token Shift and Selection Transformer for Text-Video Retrieval

    Yuqi Liu, Pengfei Xiong, Luhui Xu +2

    cs.CVarXiv:2207.07852v12022
  11. Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

    NVIDIA, :, Alisson Azzolini +51

    cs.AIcs.CVcs.LGarXiv:2503.15558v32025
  12. Content-Aware Unsupervised Deep Homography Estimation

    Jirong Zhang, Chuan Wang, Shuaicheng Liu +5

    cs.CVarXiv:1909.05983v22019
  13. Multi-Adapter RGBT Tracking

    Chenglong Li, Andong Lu, Aihua Zheng +2

    cs.CVarXiv:1907.07485v12019
  14. Illiterate DALL-E Learns to Compose

    Gautam Singh, Fei Deng, Sungjin Ahn

    cs.CVcs.LGarXiv:2110.11405v32021
  15. Making Better Mistakes: Leveraging Class Hierarchies with Deep Networks

    Luca Bertinetto, Romain Mueller, Konstantinos Tertikas +2

    cs.CVcs.LGarXiv:1912.09393v22019
  16. PromptIR: Prompting for All-in-One Blind Image Restoration

    Vaishnav Potlapalli, Syed Waqas Zamir, Salman Khan +1

    cs.CVarXiv:2306.13090v12023
  17. WorldMem: Long-term Consistent World Simulation with Memory

    Zeqi Xiao, Yushi Lan, Yifan Zhou +4

    cs.CVarXiv:2504.12369v32025
  18. Repeatability of Multiparametric Prostate MRI Radiomics Features

    Michael Schwier, Joost van Griethuysen, Mark G Vangel +7

    cs.CVeess.IVarXiv:1807.06089v22018
  19. MeshSplatBench: A Unified Benchmark for Triangle-Based Neural Rendering

    Kaixuan Zhang, Minxian Li, Mingwu Ren +1

    cs.GRcs.CVarXiv:2609.01306v12026
  20. DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control

    Junjie Wen, Yichen Zhu, Jinming Li +3

    cs.ROcs.CVarXiv:2502.05855v32025
  21. Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions

    Jing Gu, Eliana Stefani, Qi Wu +2

    cs.CVcs.AIcs.CLarXiv:2203.12667v32022
  22. DGNet: Dual-knowledge Guided Network for Infrared Small Target Detection

    Chenglong Yu, Mingzhu Xu, Jing Wang +3

    cs.CVarXiv:2609.00666v12026
  23. Bridging Language and Spherical Space: Object-Centric Control for Text-to-Panorama Generation

    Derui Li, Qian Qiao, Yuhao Sun +2

    cs.CVarXiv:2608.20691v12026
  24. Path2ST: Hierarchical Cell-Tissue Grounded Cross-Modal Translation for Spatial Transcriptomics

    Ruochen Liu, Wei Lou

    cs.CVcs.AIcs.CLarXiv:2608.14710v12026
  25. CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models

    Mukhtiar Ali, Harsh Dubey, Sugam Mishra +1

    cs.CVcs.IRcs.MMarXiv:2608.04302v12026
  26. G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection

    Yechan Kim, JongHyun Park, Dongho Yoon +2

    cs.CVcs.AIarXiv:2607.19942v22026
  27. Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

    Rahul Sajnani, Yulia Gryaditskaya, Radomír Měch +2

    cs.CVcs.AIcs.GRarXiv:2607.19344v12026
  28. Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget

    Guoxuan Chen, Chufeng Xiao, Haoran Yang +30

    cs.CVcs.AIarXiv:2607.13125v22026
  29. Confidence-Aware Tool Orchestration for Robust Video Understanding

    Yangfan He, Yujin Choi, Jaehong Yoon

    cs.CVcs.AIarXiv:2606.26904v12026
  30. Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI

    Kairos Team, Fei Wang, Shan You +21

    cs.AIcs.CVarXiv:2606.16533v32026
  31. MeshFlow: Mesh Generation with Equivariant Flow Matching

    Qi Sun, Kiyohiro Nakayama, Jing Nathan Yan +6

    cs.GRcs.CVarXiv:2606.23489v12026
  32. Can Generalist Agents Automate Data Curation?

    Feiyang Kang, Hanze Li, Adam Nguyen +5

    cs.AIcs.CLcs.CVarXiv:2606.04261v12026
  33. Decoupled Residual Denoising Diffusion Models for Unified and Data Efficient Image-to-Image Translation

    Ziyue Lin, Jiahe Hou, Hongyu Xia +6

    cs.CVarXiv:2606.01048v12026
  34. Good Token Hunting: A Hitchhiker's Guide to Token Selection for Visual Geometry Transformers

    Shuhong Zheng, Michael Oechsle, Erik Sandström +3

    cs.CVcs.AIcs.GRarXiv:2605.23892v12026
  35. Generative Modeling with Orbit-Space Particle Flow Matching

    Sinan Wang, Jinjin He, Shenyifan Lu +3

    cs.GRcs.CVarXiv:2605.02222v12026
  36. Micro-Defects Expose Macro-Fakes: Detecting AI-Generated Images via Local Distributional Shifts

    Boxuan Zhang, Jianing Zhu, Qifan Wang +2

    cs.CVcs.AIcs.LGarXiv:2605.09296v12026
  37. HL-OutPaint: Coarse-to-Fine Video Outpainting for High-Resolution Long-Range Videos

    Jeongeun Park, Janghyeok Han, Geonung Kim +4

    cs.CVcs.GRarXiv:2605.17543v32026
  38. CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization

    Ahmed Heakl, Abdelrahman M. Shaker, Youssef Mohamed +4

    cs.LGcs.CLcs.CVarXiv:2605.19436v12026
  39. Fruit Detection, Segmentation and 3D Visualisation of Environments in Apple Orchards

    Hanwen Kang, Chao Chen

    cs.CVeess.IVarXiv:1911.12889v12019
  40. UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors

    Houyuan Chen, Hong Li, Xianghao Kong +8

    cs.CVarXiv:2605.00658v12026
  41. RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics

    Enshen Zhou, Jingkun An, Cheng Chi +8

    cs.ROcs.AIcs.CVarXiv:2506.04308v42025
  42. Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance

    Jason Qiu, Zachary Meurer, Xavier Thomas +1

    cs.CVarXiv:2604.01848v42026
  43. On-the-fly Repulsion in the Contextual Space for Rich Diversity in Diffusion Transformers

    Omer Dahary, Benaya Koren, Daniel Garibi +1

    cs.CVcs.AIcs.GRarXiv:2603.28762v22026
  44. ExpArt-KG: Artwork Image Description Generation through Iterative Exploration of Knowledge Graphs

    Yuta Kato, Shintaro Ozaki, Kazuki Hayashi +4

    cs.CLcs.CVarXiv:2609.00629v12026
  45. Less Gaussians, Texture More: 4K Feed-Forward Textured Splatting

    Yixing Lao, Xuyang Bai, Xiaoyang Wu +7

    cs.CVarXiv:2603.25745v12026
  46. Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding

    Yinghui Li, Jiayi Kuang, Peng Xing +11

    cs.AIcs.CVarXiv:2603.18472v22026
  47. Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis

    Tianbao Xie, Jiaqi Deng, Xiaochuan Li +12

    cs.AIcs.CLcs.CVarXiv:2505.13227v32025
  48. GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing

    Mingxin Liu, Ziqian Fan, Zhaokai Wang +13

    cs.CVarXiv:2603.12264v12026
  49. DanceFormer: Music Conditioned 3D Dance Generation with Parametric Motion Transformer

    Buyu Li, Yongchi Zhao, Zhelun Shi +1

    cs.AIcs.CVarXiv:2103.10206v52021
  50. Memory Attention Networks for Skeleton-based Action Recognition

    Chunyu Xie, Ce Li, Baochang Zhang +4

    cs.CVarXiv:1804.08254v22018
  51. Towards Practical Real-Time Neural Video Compression

    Zhaoyang Jia, Bin Li, Jiahao Li +4

    eess.IVcs.CVarXiv:2502.20762v22025
  52. LongCat-Image Technical Report

    Meituan LongCat Team, Hanghang Ma, Haoxian Tan +10

    cs.CVarXiv:2512.07584v12025
  53. LongVPO: From Anchored Cues to Self-Reasoning for Long-Form Video Preference Optimization

    Zhenpeng Huang, Jiaqi Li, Zihan Jia +6

    cs.CVarXiv:2602.02341v12026
  54. Think-Then-Generate: Reasoning-Aware Text-to-Image Diffusion with LLM Encoders

    Siqi Kou, Jiachun Jin, Zetong Zhou +8

    cs.CVarXiv:2601.10332v12026
  55. History-Guided Video Diffusion

    Kiwhan Song, Boyuan Chen, Max Simchowitz +3

    cs.LGcs.CVarXiv:2502.06764v22025
  56. Deep Learning with Lung Segmentation and Bone Shadow Exclusion Techniques for Chest X-Ray Analysis of Lung Cancer

    Yu. Gordienko, Peng Gang, Jiang Hui +5

    cs.LGcs.CVarXiv:1712.07632v12017
  57. RealNet: A Feature Selection Network with Realistic Synthetic Anomaly for Anomaly Detection

    Ximiao Zhang, Min Xu, Xiuzhuang Zhou

    cs.CVarXiv:2403.05897v12024
  58. Gaussian Splatting SLAM

    Hidenobu Matsuki, Riku Murai, Paul H. J. Kelly +1

    cs.CVcs.ROarXiv:2312.06741v22023
  59. Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification

    Aojun Zhou, Ke Wang, Zimu Lu +8

    cs.CLcs.AIcs.CVarXiv:2308.07921v12023
  60. ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation

    Jiazheng Xu, Xiao Liu, Yuchen Wu +5

    cs.CVcs.LGarXiv:2304.05977v42023