Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

361 to 420 of 18,815

  1. Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation

    Pengxiang Li, Zechen Hu, Zirui Shang +15

    cs.LGcs.AIcs.CVarXiv:2509.23866v12025
  2. Evolving Losses for Unsupervised Video Representation Learning

    AJ Piergiovanni, Anelia Angelova, Michael S. Ryoo

    cs.CVcs.LGarXiv:2002.12177v12020
  3. LIBRE: The Multiple 3D LiDAR Dataset

    Alexander Carballo, Jacob Lambert, Abraham Monrroy-Cano +6

    cs.ROcs.CVarXiv:2003.06129v22020
  4. Parallel Attention: A Unified Framework for Visual Object Discovery through Dialogs and Queries

    Bohan Zhuang, Qi Wu, Chunhua Shen +2

    cs.CVarXiv:1711.06370v12017
  5. NeuroSymbEAD: A Large Scale Neuro-Symbolic Caption Dataset for Omni-Directional Embodied Autonomous Driving

    Muhammad Ahmed Ullah Khan, Mohammed Elamine, Sheikh Talha Uddin +3

    cs.CVarXiv:2609.16919v12026
  6. Global Context Networks

    Yue Cao, Jiarui Xu, Stephen Lin +2

    cs.CVarXiv:2012.13375v12020
  7. LM-PCVMNet: Pediatric Cervical Vertebral Maturation Analysis with Deep Fusion of Landmarks and Metadata

    Peng Wang, Wanzhen Song, Anli Wang +3

    eess.IVcs.CVarXiv:2609.16033v12026
  8. Towards Blind Watermarking: Combining Invertible and Non-invertible Mechanisms

    Rui Ma, Mengxi Guo, Yi Hou +4

    cs.MMcs.CVarXiv:2212.12678v12022
  9. Data-Free Network Quantization With Adversarial Knowledge Distillation

    Yoojin Choi, Jihwan Choi, Mostafa El-Khamy +1

    cs.CVcs.LGcs.NEarXiv:2005.04136v12020
  10. WebQA: Multihop and Multimodal QA

    Yingshan Chang, Mridu Narang, Hisami Suzuki +3

    cs.CLcs.AIcs.CVarXiv:2109.00590v42021
  11. Convolution-Free Medical Image Segmentation using Transformers

    Davood Karimi, Serge Vasylechko, Ali Gholipour

    eess.IVcs.CVarXiv:2102.13645v22021
  12. Evaluating Mesh Reconstruction Methods for Crop Phenotyping

    Karanvir Singh, Theo Morales, Binh-Son Hua +1

    cs.CVarXiv:2609.16926v12026
  13. Augmentation Matters: A Simple-yet-Effective Approach to Semi-supervised Semantic Segmentation

    Zhen Zhao, Lihe Yang, Sifan Long +3

    cs.CVarXiv:2212.04976v12022
  14. From Foundation Embeddings to Cropland Maps: Label Efficiency, Temporal Transferability and Independent Human Validation

    Mohammad Ammar Mughees, Giovanni Montefoschi, Zhongxin Chen +1

    cs.CVcs.LGeess.IVarXiv:2609.17138v12026
  15. MAETrack: Unleashing the Potential of Pretrained Geometric Priors for 3D Single Object Tracking

    Sifan Zhou, Qiwei Wang, Linyue Tan +3

    cs.CVarXiv:2609.16695v12026
  16. A Comprehensive Review of Deep Learning-based Single Image Super-resolution

    Syed Muhammad Arsalan Bashir, Yi Wang, Mahrukh Khan +1

    cs.CVcs.LGeess.IVarXiv:2102.09351v32021
  17. SAVTrack: Selective Vote Aggregation for Reliability-Aware Point Cloud Tracking

    Sifan Zhou, Linyue Tan, Qiwei Wang +2

    cs.CVarXiv:2609.16662v12026
  18. Vision-based Human Fall Detection Systems using Deep Learning: A Review

    Ekram Alam, Abu Sufian, Paramartha Dutta +1

    cs.CVcs.AIarXiv:2207.10952v12022
  19. Boosting Few-shot Fine-grained Recognition with Background Suppression and Foreground Alignment

    Zican Zha, Hao Tang, Yunlian Sun +1

    cs.CVarXiv:2210.01439v22022
  20. Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications

    Han Cai, Ji Lin, Yujun Lin +5

    cs.LGcs.CLcs.CVarXiv:2204.11786v12022
  21. Monocular Depth Estimation: A Survey

    Amlaan Bhoi

    cs.CVarXiv:1901.09402v12019
  22. A Multi-Vehicle Dataset with Camera, LiDAR, and Radar Sensors and Scanned 3D Models for Custom Auto-Annotation using RTK-GNSS

    Philipp Berthold, Bianca Forkel, Mirko Maehlisch

    cs.ROcs.AIcs.CVarXiv:2609.12871v12026
  23. HyperTransformer: A Textural and Spectral Feature Fusion Transformer for Pansharpening

    Wele Gedara Chaminda Bandara, Vishal M. Patel

    cs.CVeess.IVarXiv:2203.02503v32022
  24. Differentiable Mesh State Estimation via Factor Graph Inference for Deformable Object Reconstruction

    Lidia Al-Zogbi, Fangjie Li, Samuel Tobin +13

    cs.ROcs.CVarXiv:2609.16686v12026
  25. Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance

    Yujie Wei, Shiwei Zhang, Hangjie Yuan +8

    cs.CVarXiv:2510.24711v22025
  26. ResLRP: The Role of Residual Cancellation in Attribution Instability in Vision Transformers

    Jim Berend, Reduan Achtibat, Daniel Schäffer +4

    cs.CVcs.AIcs.LGarXiv:2609.17152v12026
  27. Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

    John Won, Kyungmin Lee, Huiwon Jang +2

    cs.CVcs.ROarXiv:2510.27607v32025
  28. LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models

    Hao Zhang, Hongyang Li, Feng Li +8

    cs.CVarXiv:2312.02949v12023
  29. DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models

    Zefeng He, Xiaoye Qu, Yafu Li +3

    cs.CVarXiv:2512.24165v12025
  30. Unsupervised Semantic Correspondence Using Stable Diffusion

    Eric Hedlin, Gopal Sharma, Shweta Mahajan +4

    cs.CVarXiv:2305.15581v22023
  31. Spectral-Spatial Mamba for Hyperspectral Image Classification

    Lingbo Huang, Yushi Chen, Xin He

    cs.CVarXiv:2404.18401v32024
  32. Which Pretext Task Transfers? Self-Supervised Pretraining Objectives for Lung Ultrasound

    Moein Heidari, Junbo Rao, Jai Choraria +3

    cs.CVarXiv:2609.16551v12026
  33. Vision And Text Transformer For Predicting Answerability On Visual Question Answering

    Tung Le, Huy Tien Nguyen, Le Minh Nguyen

    cs.CVcs.AIarXiv:2609.16565v12026
  34. SSC-Priors: Exploring Semantic and Visibility Priors to Boost Lidar Semantic Scene Completion

    Tetiana Martyniuk, Jonathan Seele, Alexandre Boulch +3

    cs.CVarXiv:2609.17413v12026
  35. gradSim: Differentiable simulation for system identification and visuomotor control

    Krishna Murthy Jatavallabhula, Miles Macklin, Florian Golemo +11

    cs.CVcs.AIcs.LGarXiv:2104.02646v12021
  36. Neuro-Symbolic Hierarchical Intention Anticipation in Human Behavior

    Farnaz Soleimani, Abdelghani Chibani, Yacine Amirat +1

    cs.AIcs.CVcs.HCarXiv:2609.17064v12026
  37. MANTRA: Memory Augmented Networks for Multiple Trajectory Prediction

    Francesco Marchetti, Federico Becattini, Lorenzo Seidenari +1

    cs.CVarXiv:2006.03340v22020
  38. TEMPO: Learning Temporal Context for Dynamic Robot Manipulation

    Zhenyang Feng, Jimin Heo, Erik B. Sudderth +1

    cs.ROcs.CVcs.LGarXiv:2609.16864v12026
  39. An Empirical Study of Language CNN for Image Captioning

    Jiuxiang Gu, Gang Wang, Jianfei Cai +1

    cs.CVcs.LGarXiv:1612.07086v32016
  40. DeVRF: Fast Deformable Voxel Radiance Fields for Dynamic Scenes

    Jia-Wei Liu, Yan-Pei Cao, Weijia Mao +6

    cs.CVarXiv:2205.15723v22022
  41. Convolutional Neural Networks for Global Human Settlements Mapping from Sentinel-2 Satellite Imagery

    Christina Corbane, Vasileios Syrris, Filip Sabo +5

    eess.IVcs.CVcs.LGarXiv:2006.03267v22020
  42. MUMINS: Metadata-conditioned Uncertainty-aware Medical Image Next-state Synthesis

    Anna Oliveras, Roger Marí, Rafael Redondo +7

    cs.CVcs.AIarXiv:2609.17169v12026
  43. GraLoD: Graphics-Inspired Continuous Level-of-Detail Learning for Image Restoration

    Hu Gao, Lizhuang Ma, Yulong Chen

    cs.CVarXiv:2609.16578v12026
  44. DetGPT: Detect What You Need via Reasoning

    Renjie Pi, Jiahui Gao, Shizhe Diao +8

    cs.CVcs.AIarXiv:2305.14167v22023
  45. Seeing What Matters: Visual Cue Guided Video Planning for Generalizable Robot Navigation

    Hojin Lee, Sizhe Lester Li, Maximilian Hilger +4

    cs.ROcs.AIcs.CVarXiv:2609.16737v12026
  46. DexNDM: Closing the Reality Gap for Dexterous In-Hand Rotation via Joint-Wise Neural Dynamics Model

    Xueyi Liu, He Wang, Li Yi

    cs.ROcs.CVarXiv:2510.08556v12025
  47. Accelerated Decoding of Centroid Positional Encoding for Instance Segmentation

    Carmelo Scribano, Filippo Muzzini, Nedyalko Prisadnikov +6

    cs.CVarXiv:2609.16874v12026
  48. VChain: Chain-of-Visual-Thought for Reasoning in Video Generation

    Ziqi Huang, Ning Yu, Gordon Chen +3

    cs.CVarXiv:2510.05094v22025
  49. GeoSVR: Taming Sparse Voxels for Geometrically Accurate Surface Reconstruction

    Jiahe Li, Jiawei Zhang, Youmin Zhang +4

    cs.CVarXiv:2509.18090v22025
  50. Deep Convolutional Neural Networks for Interpretable Analysis of EEG Sleep Stage Scoring

    Albert Vilamala, Kristoffer H. Madsen, Lars K. Hansen

    cs.CVstat.MLarXiv:1710.00633v12017
  51. Deep Learning Based Steel Pipe Weld Defect Detection

    Dingming Yang, Yanrong Cui, Zeyu Yu +1

    cs.CVcs.AIarXiv:2104.14907v22021
  52. Motion Mamba: Efficient and Long Sequence Motion Generation

    Zeyu Zhang, Akide Liu, Ian Reid +3

    cs.CVarXiv:2403.07487v42024
  53. PTB-TIR: A Thermal Infrared Pedestrian Tracking Benchmark

    Qiao Liu, Zhenyu He, Xin Li +1

    cs.CVarXiv:1801.05944v32018
  54. Jointly Cross- and Self-Modal Graph Attention Network for Query-Based Moment Localization

    Daizong Liu, Xiaoye Qu, Xiao-Yang Liu +3

    cs.CVcs.IRarXiv:2008.01403v22020
  55. FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation

    Guangyu Sun, Shlok Kumar Mishra, Wentao Bao +8

    cs.CVarXiv:2609.16591v12026
  56. GraspNeRF: Multiview-based 6-DoF Grasp Detection for Transparent and Specular Objects Using Generalizable NeRF

    Qiyu Dai, Yan Zhu, Yiran Geng +3

    cs.ROcs.CVarXiv:2210.06575v32022
  57. AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video

    Jiaming Tan, Mingliang Zhai, Zhen Li +3

    cs.CVarXiv:2609.14462v12026
  58. Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation

    Niantong Li, Guangzheng Hu, Weixu Qiao +35

    cs.CVarXiv:2605.28091v22026
  59. Frequency Perception Network for Camouflaged Object Detection

    Runmin Cong, Mengyao Sun, Sanyi Zhang +3

    cs.CVarXiv:2308.08924v22023
  60. Towards Zero-Shot Scale-Aware Monocular Depth Estimation

    Vitor Guizilini, Igor Vasiljevic, Dian Chen +2

    cs.CVcs.LGarXiv:2306.17253v12023