Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
361 to 420 of 18,815
Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
Pengxiang Li, Zechen Hu, Zirui Shang +15
cs.LGcs.AIcs.CVarXiv:2509.23866v12025Evolving Losses for Unsupervised Video Representation Learning
AJ Piergiovanni, Anelia Angelova, Michael S. Ryoo
cs.CVcs.LGarXiv:2002.12177v12020LIBRE: The Multiple 3D LiDAR Dataset
Alexander Carballo, Jacob Lambert, Abraham Monrroy-Cano +6
cs.ROcs.CVarXiv:2003.06129v22020Parallel Attention: A Unified Framework for Visual Object Discovery through Dialogs and Queries
Bohan Zhuang, Qi Wu, Chunhua Shen +2
cs.CVarXiv:1711.06370v12017NeuroSymbEAD: A Large Scale Neuro-Symbolic Caption Dataset for Omni-Directional Embodied Autonomous Driving
Muhammad Ahmed Ullah Khan, Mohammed Elamine, Sheikh Talha Uddin +3
cs.CVarXiv:2609.16919v12026Global Context Networks
Yue Cao, Jiarui Xu, Stephen Lin +2
cs.CVarXiv:2012.13375v12020LM-PCVMNet: Pediatric Cervical Vertebral Maturation Analysis with Deep Fusion of Landmarks and Metadata
Peng Wang, Wanzhen Song, Anli Wang +3
eess.IVcs.CVarXiv:2609.16033v12026Towards Blind Watermarking: Combining Invertible and Non-invertible Mechanisms
Rui Ma, Mengxi Guo, Yi Hou +4
cs.MMcs.CVarXiv:2212.12678v12022Data-Free Network Quantization With Adversarial Knowledge Distillation
Yoojin Choi, Jihwan Choi, Mostafa El-Khamy +1
cs.CVcs.LGcs.NEarXiv:2005.04136v12020WebQA: Multihop and Multimodal QA
Yingshan Chang, Mridu Narang, Hisami Suzuki +3
cs.CLcs.AIcs.CVarXiv:2109.00590v42021Convolution-Free Medical Image Segmentation using Transformers
Davood Karimi, Serge Vasylechko, Ali Gholipour
eess.IVcs.CVarXiv:2102.13645v22021Evaluating Mesh Reconstruction Methods for Crop Phenotyping
Karanvir Singh, Theo Morales, Binh-Son Hua +1
cs.CVarXiv:2609.16926v12026Augmentation Matters: A Simple-yet-Effective Approach to Semi-supervised Semantic Segmentation
Zhen Zhao, Lihe Yang, Sifan Long +3
cs.CVarXiv:2212.04976v12022From Foundation Embeddings to Cropland Maps: Label Efficiency, Temporal Transferability and Independent Human Validation
Mohammad Ammar Mughees, Giovanni Montefoschi, Zhongxin Chen +1
cs.CVcs.LGeess.IVarXiv:2609.17138v12026MAETrack: Unleashing the Potential of Pretrained Geometric Priors for 3D Single Object Tracking
Sifan Zhou, Qiwei Wang, Linyue Tan +3
cs.CVarXiv:2609.16695v12026A Comprehensive Review of Deep Learning-based Single Image Super-resolution
Syed Muhammad Arsalan Bashir, Yi Wang, Mahrukh Khan +1
cs.CVcs.LGeess.IVarXiv:2102.09351v32021SAVTrack: Selective Vote Aggregation for Reliability-Aware Point Cloud Tracking
Sifan Zhou, Linyue Tan, Qiwei Wang +2
cs.CVarXiv:2609.16662v12026Vision-based Human Fall Detection Systems using Deep Learning: A Review
Ekram Alam, Abu Sufian, Paramartha Dutta +1
cs.CVcs.AIarXiv:2207.10952v12022Boosting Few-shot Fine-grained Recognition with Background Suppression and Foreground Alignment
Zican Zha, Hao Tang, Yunlian Sun +1
cs.CVarXiv:2210.01439v22022Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
Han Cai, Ji Lin, Yujun Lin +5
cs.LGcs.CLcs.CVarXiv:2204.11786v12022Monocular Depth Estimation: A Survey
Amlaan Bhoi
cs.CVarXiv:1901.09402v12019A Multi-Vehicle Dataset with Camera, LiDAR, and Radar Sensors and Scanned 3D Models for Custom Auto-Annotation using RTK-GNSS
Philipp Berthold, Bianca Forkel, Mirko Maehlisch
cs.ROcs.AIcs.CVarXiv:2609.12871v12026HyperTransformer: A Textural and Spectral Feature Fusion Transformer for Pansharpening
Wele Gedara Chaminda Bandara, Vishal M. Patel
cs.CVeess.IVarXiv:2203.02503v32022Differentiable Mesh State Estimation via Factor Graph Inference for Deformable Object Reconstruction
Lidia Al-Zogbi, Fangjie Li, Samuel Tobin +13
cs.ROcs.CVarXiv:2609.16686v12026Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance
Yujie Wei, Shiwei Zhang, Hangjie Yuan +8
cs.CVarXiv:2510.24711v22025ResLRP: The Role of Residual Cancellation in Attribution Instability in Vision Transformers
Jim Berend, Reduan Achtibat, Daniel Schäffer +4
cs.CVcs.AIcs.LGarXiv:2609.17152v12026Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
John Won, Kyungmin Lee, Huiwon Jang +2
cs.CVcs.ROarXiv:2510.27607v32025LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models
Hao Zhang, Hongyang Li, Feng Li +8
cs.CVarXiv:2312.02949v12023DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models
Zefeng He, Xiaoye Qu, Yafu Li +3
cs.CVarXiv:2512.24165v12025Unsupervised Semantic Correspondence Using Stable Diffusion
Eric Hedlin, Gopal Sharma, Shweta Mahajan +4
cs.CVarXiv:2305.15581v22023Spectral-Spatial Mamba for Hyperspectral Image Classification
Lingbo Huang, Yushi Chen, Xin He
cs.CVarXiv:2404.18401v32024Which Pretext Task Transfers? Self-Supervised Pretraining Objectives for Lung Ultrasound
Moein Heidari, Junbo Rao, Jai Choraria +3
cs.CVarXiv:2609.16551v12026Vision And Text Transformer For Predicting Answerability On Visual Question Answering
Tung Le, Huy Tien Nguyen, Le Minh Nguyen
cs.CVcs.AIarXiv:2609.16565v12026SSC-Priors: Exploring Semantic and Visibility Priors to Boost Lidar Semantic Scene Completion
Tetiana Martyniuk, Jonathan Seele, Alexandre Boulch +3
cs.CVarXiv:2609.17413v12026gradSim: Differentiable simulation for system identification and visuomotor control
Krishna Murthy Jatavallabhula, Miles Macklin, Florian Golemo +11
cs.CVcs.AIcs.LGarXiv:2104.02646v12021Neuro-Symbolic Hierarchical Intention Anticipation in Human Behavior
Farnaz Soleimani, Abdelghani Chibani, Yacine Amirat +1
cs.AIcs.CVcs.HCarXiv:2609.17064v12026MANTRA: Memory Augmented Networks for Multiple Trajectory Prediction
Francesco Marchetti, Federico Becattini, Lorenzo Seidenari +1
cs.CVarXiv:2006.03340v22020TEMPO: Learning Temporal Context for Dynamic Robot Manipulation
Zhenyang Feng, Jimin Heo, Erik B. Sudderth +1
cs.ROcs.CVcs.LGarXiv:2609.16864v12026An Empirical Study of Language CNN for Image Captioning
Jiuxiang Gu, Gang Wang, Jianfei Cai +1
cs.CVcs.LGarXiv:1612.07086v32016DeVRF: Fast Deformable Voxel Radiance Fields for Dynamic Scenes
Jia-Wei Liu, Yan-Pei Cao, Weijia Mao +6
cs.CVarXiv:2205.15723v22022Convolutional Neural Networks for Global Human Settlements Mapping from Sentinel-2 Satellite Imagery
Christina Corbane, Vasileios Syrris, Filip Sabo +5
eess.IVcs.CVcs.LGarXiv:2006.03267v22020MUMINS: Metadata-conditioned Uncertainty-aware Medical Image Next-state Synthesis
Anna Oliveras, Roger Marí, Rafael Redondo +7
cs.CVcs.AIarXiv:2609.17169v12026GraLoD: Graphics-Inspired Continuous Level-of-Detail Learning for Image Restoration
Hu Gao, Lizhuang Ma, Yulong Chen
cs.CVarXiv:2609.16578v12026DetGPT: Detect What You Need via Reasoning
Renjie Pi, Jiahui Gao, Shizhe Diao +8
cs.CVcs.AIarXiv:2305.14167v22023Seeing What Matters: Visual Cue Guided Video Planning for Generalizable Robot Navigation
Hojin Lee, Sizhe Lester Li, Maximilian Hilger +4
cs.ROcs.AIcs.CVarXiv:2609.16737v12026DexNDM: Closing the Reality Gap for Dexterous In-Hand Rotation via Joint-Wise Neural Dynamics Model
Xueyi Liu, He Wang, Li Yi
cs.ROcs.CVarXiv:2510.08556v12025Accelerated Decoding of Centroid Positional Encoding for Instance Segmentation
Carmelo Scribano, Filippo Muzzini, Nedyalko Prisadnikov +6
cs.CVarXiv:2609.16874v12026VChain: Chain-of-Visual-Thought for Reasoning in Video Generation
Ziqi Huang, Ning Yu, Gordon Chen +3
cs.CVarXiv:2510.05094v22025GeoSVR: Taming Sparse Voxels for Geometrically Accurate Surface Reconstruction
Jiahe Li, Jiawei Zhang, Youmin Zhang +4
cs.CVarXiv:2509.18090v22025Deep Convolutional Neural Networks for Interpretable Analysis of EEG Sleep Stage Scoring
Albert Vilamala, Kristoffer H. Madsen, Lars K. Hansen
cs.CVstat.MLarXiv:1710.00633v12017Deep Learning Based Steel Pipe Weld Defect Detection
Dingming Yang, Yanrong Cui, Zeyu Yu +1
cs.CVcs.AIarXiv:2104.14907v22021Motion Mamba: Efficient and Long Sequence Motion Generation
Zeyu Zhang, Akide Liu, Ian Reid +3
cs.CVarXiv:2403.07487v42024PTB-TIR: A Thermal Infrared Pedestrian Tracking Benchmark
Qiao Liu, Zhenyu He, Xin Li +1
cs.CVarXiv:1801.05944v32018Jointly Cross- and Self-Modal Graph Attention Network for Query-Based Moment Localization
Daizong Liu, Xiaoye Qu, Xiao-Yang Liu +3
cs.CVcs.IRarXiv:2008.01403v22020FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation
Guangyu Sun, Shlok Kumar Mishra, Wentao Bao +8
cs.CVarXiv:2609.16591v12026GraspNeRF: Multiview-based 6-DoF Grasp Detection for Transparent and Specular Objects Using Generalizable NeRF
Qiyu Dai, Yan Zhu, Yiran Geng +3
cs.ROcs.CVarXiv:2210.06575v32022AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video
Jiaming Tan, Mingliang Zhai, Zhen Li +3
cs.CVarXiv:2609.14462v12026Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation
Niantong Li, Guangzheng Hu, Weixu Qiao +35
cs.CVarXiv:2605.28091v22026Frequency Perception Network for Camouflaged Object Detection
Runmin Cong, Mengyao Sun, Sanyi Zhang +3
cs.CVarXiv:2308.08924v22023Towards Zero-Shot Scale-Aware Monocular Depth Estimation
Vitor Guizilini, Igor Vasiljevic, Dian Chen +2
cs.CVcs.LGarXiv:2306.17253v12023