Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

3,541 to 3,600 of 18,822

  1. FocalClick: Towards Practical Interactive Image Segmentation

    Xi Chen, Zhiyan Zhao, Yilei Zhang +3

    cs.CVarXiv:2204.02574v22022
  2. RAM: Recover Any 3D Human Motion in-the-Wild

    Sen Jia, Ning Zhu, Jinqin Zhong +4

    cs.CVcs.AIarXiv:2603.19929v22026
  3. The Paradigm Shift: A Comprehensive Survey on Large Vision Language Models for Multimodal Fake News Detection

    Wei Ai, Yilong Tan, Yuntao Shou +4

    cs.AIcs.CVarXiv:2601.15316v12026
  4. Rope3D: TheRoadside Perception Dataset for Autonomous Driving and Monocular 3D Object Detection Task

    Xiaoqing Ye, Mao Shu, Hanyu Li +5

    cs.CVarXiv:2203.13608v12022
  5. Visual News: Benchmark and Challenges in News Image Captioning

    Fuxiao Liu, Yinghan Wang, Tianlu Wang +1

    cs.CVarXiv:2010.03743v32020
  6. AdvPC: Transferable Adversarial Perturbations on 3D Point Clouds

    Abdullah Hamdi, Sara Rojas, Ali Thabet +1

    cs.CVcs.CRcs.LGarXiv:1912.00461v22019
  7. Image-to-Lidar Self-Supervised Distillation for Autonomous Driving Data

    Corentin Sautier, Gilles Puy, Spyros Gidaris +3

    cs.CVcs.LGarXiv:2203.16258v12022
  8. UNICON: Combating Label Noise Through Uniform Selection and Contrastive Learning

    Nazmul Karim, Mamshad Nayeem Rizve, Nazanin Rahnavard +2

    cs.CVcs.LGarXiv:2203.14542v42022
  9. Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning

    Hulingxiao He, Zijun Geng, Yuxin Peng

    cs.CVcs.AIarXiv:2602.07605v32026
  10. B-CNN: Branch Convolutional Neural Network for Hierarchical Classification

    Xinqi Zhu, Michael Bain

    cs.CVarXiv:1709.09890v22017
  11. Learning to score the figure skating sports videos

    Chengming Xu, Yanwei Fu, Bing Zhang +3

    cs.MMcs.CVarXiv:1802.02774v32018
  12. Nighttime Dehazing with a Synthetic Benchmark

    Jing Zhang, Yang Cao, Zheng-Jun Zha +1

    cs.CVcs.LGeess.IVarXiv:2008.03864v32020
  13. Rethinking Data Augmentation for Image Super-resolution: A Comprehensive Analysis and a New Strategy

    Jaejun Yoo, Namhyuk Ahn, Kyung-Ah Sohn

    eess.IVcs.CVarXiv:2004.00448v22020
  14. Cross-Domain Few-Shot Classification via Adversarial Task Augmentation

    Haoqing Wang, Zhi-Hong Deng

    cs.CVarXiv:2104.14385v22021
  15. Simple Unsupervised Object-Centric Learning for Complex and Naturalistic Videos

    Gautam Singh, Yi-Fu Wu, Sungjin Ahn

    cs.CVcs.LGarXiv:2205.14065v12022
  16. Learning to Hash with Binary Deep Neural Network

    Thanh-Toan Do, Anh-Dzung Doan, Ngai-Man Cheung

    cs.CVarXiv:1607.05140v12016
  17. MCPO: Modality-Contrastive Preference Optimization for Multimodal Chain-of-Thought Compression

    Guangheng Yang, Zhenliang Ni, Zhenkai Wu +4

    cs.CVcs.AIarXiv:2609.04947v12026
  18. Language-Conditioned World Modeling for Visual Navigation

    Yifei Dong, Fengyi Wu, Yilong Dai +10

    cs.CVcs.AIcs.ROarXiv:2603.26741v12026
  19. StableWorld: Towards Stable and Consistent Long Interactive Video Generation

    Ying Yang, Zhengyao Lv, Yujia Zeng +9

    cs.CVarXiv:2601.15281v22026
  20. One Diffusion Model, Two Roles: Guided Trajectory Planning and Safety-Critical Scenario Generation in Closed-Loop Simulation

    Arka Pal, Rajesh Kumar, Hannes Eriksson +4

    cs.CVcs.AIcs.LGarXiv:2609.04921v12026
  21. Efficient Test-Time Adaptation of Vision-Language Models

    Adilbek Karmanov, Dayan Guan, Shijian Lu +2

    cs.CVarXiv:2403.18293v12024
  22. Sound-based Multi-Person 3D Pose Estimation

    Yusuke Oumi, Yuto Shibata, Go Irie +3

    cs.CVcs.AIcs.LGarXiv:2609.04902v12026
  23. The reliability of a deep learning model in clinical out-of-distribution MRI data: a multicohort study

    Gustav Mårtensson, Daniel Ferreira, Tobias Granberg +22

    physics.med-phcs.CVcs.LGarXiv:1911.00515v12019
  24. Unravelling Robustness of Deep Learning based Face Recognition Against Adversarial Attacks

    Gaurav Goswami, Nalini Ratha, Akshay Agarwal +2

    cs.CVarXiv:1803.00401v12018
  25. Highly Accurate Dichotomous Image Segmentation

    Xuebin Qin, Hang Dai, Xiaobin Hu +3

    cs.CVarXiv:2203.03041v42022
  26. Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form Sentences

    Zhu Zhang, Zhou Zhao, Yang Zhao +3

    cs.CVarXiv:2001.06891v32020
  27. LVSM: A Large View Synthesis Model with Minimal 3D Inductive Bias

    Haian Jin, Hanwen Jiang, Hao Tan +6

    cs.CVcs.GRcs.LGarXiv:2410.17242v22024
  28. Parametric Classification for Generalized Category Discovery: A Baseline Study

    Xin Wen, Bingchen Zhao, Xiaojuan Qi

    cs.CVcs.LGarXiv:2211.11727v42022
  29. Methane Detection On Board Satellites from Unorthorectified Imagery

    Luca Marini, Maggie Chen, Hala Lamdouar +3

    cs.CVcs.AIcs.LGarXiv:2609.04906v12026
  30. A very preliminary analysis of DALL-E 2

    Gary Marcus, Ernest Davis, Scott Aaronson

    cs.CVcs.AIarXiv:2204.13807v22022
  31. SimFuse3D: Source-Guided Target Simulation and Confidence-Guided Multi-Stage Localization Reweighting for Cross-Platform 3D Object Detection

    Yongchun Lin, Xinliang Zhang, Yun Zou +7

    cs.CVcs.AIarXiv:2609.04886v12026
  32. Foveation-based Mechanisms Alleviate Adversarial Examples

    Yan Luo, Xavier Boix, Gemma Roig +2

    cs.LGcs.CVarXiv:1511.06292v32015
  33. Mitigating Performance Discrepancy in Cross-Domain 3D Class-Incremental Learning

    Jinge Ma, Gautham Vinod, Bruce Coburn +3

    cs.CVcs.AIarXiv:2609.04860v12026
  34. Adversarial Example Detection for DNN Models: A Review and Experimental Comparison

    Ahmed Aldahdooh, Wassim Hamidouche, Sid Ahmed Fezza +1

    cs.CVcs.CRarXiv:2105.00203v42021
  35. VPFNet: Improving 3D Object Detection with Virtual Point based LiDAR and Stereo Data Fusion

    Hanqi Zhu, Jiajun Deng, Yu Zhang +4

    cs.CVarXiv:2111.14382v22021
  36. Not All Directions Matter: Towards Structured and Task-Aware Low-Rank Model Adaptation

    Xi Xiao, Chenrui Ma, Yunbei Zhang +7

    cs.CVarXiv:2603.14228v22026
  37. SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning

    Jian Zhang, Shijie Zhou, Bangya Liu +2

    cs.CVarXiv:2603.27437v32026
  38. DealMVC: Dual Contrastive Calibration for Multi-view Clustering

    Xihong Yang, Jiaqi Jin, Siwei Wang +7

    cs.CVcs.LGarXiv:2308.09000v32023
  39. Linguistic Trajectory Encoding for Efficient Long-Horizon Spatial Memory in Embodied Agents

    Tianyidan Xie, Shenyi Wang, Qiang Tang +7

    cs.CVcs.AIarXiv:2609.04802v12026
  40. Satellite Imagery Feature Detection using Deep Convolutional Neural Network: A Kaggle Competition

    Vladimir Iglovikov, Sergey Mushinskiy, Vladimir Osin

    cs.CVarXiv:1706.06169v12017
  41. Pansharpening for Thin-Cloud Contaminated Remote Sensing Images: A Unified Framework and Benchmark Dataset

    Songcheng Du, Yang Zou, Jiaxin Li +4

    cs.CVarXiv:2603.14952v12026
  42. Dynamic High-frequency Convolution for Infrared Small Target Detection

    Ruojing Li, Chao Xiao, Qian Yin +5

    cs.CVarXiv:2602.02969v22026
  43. Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild

    Changda Zhou, Ziyue Gao, Xueqing Wang +4

    cs.CVarXiv:2603.04205v22026
  44. HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis

    Mingjin Chen, Junhao Chen, Zhaoxin Fan +6

    cs.CVarXiv:2604.03305v12026
  45. MR Image Denoising and Super-Resolution Using Regularized Reverse Diffusion

    Hyungjin Chung, Eun Sun Lee, Jong Chul Ye

    eess.IVcs.AIcs.CVarXiv:2203.12621v12022
  46. Efficient Medical Image Segmentation Based on Knowledge Distillation

    Dian Qin, Jiajun Bu, Zhe Liu +6

    eess.IVcs.CVarXiv:2108.09987v12021
  47. Unsupervised Domain Adaptation using Feature-Whitening and Consensus Loss

    Subhankar Roy, Aliaksandr Siarohin, Enver Sangineto +3

    cs.CVarXiv:1903.03215v22019
  48. DAE-Former: Dual Attention-guided Efficient Transformer for Medical Image Segmentation

    Reza Azad, René Arimond, Ehsan Khodapanah Aghdam +2

    cs.CVarXiv:2212.13504v32022
  49. SAM-CLIP: Merging Vision Foundation Models towards Semantic and Spatial Understanding

    Haoxiang Wang, Pavan Kumar Anasosalu Vasu, Fartash Faghri +6

    cs.CVcs.LGarXiv:2310.15308v42023
  50. Convolutional Kolmogorov-Arnold Networks

    Alexander Dylan Bodner, Antonio Santiago Tepsich, Jack Natan Spolski +1

    cs.CVcs.AIarXiv:2406.13155v32024
  51. Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies

    Wenjin Hou, Wei Liu, Han Hu +3

    cs.CVarXiv:2602.01816v12026
  52. RoIMix: Proposal-Fusion among Multiple Images for Underwater Object Detection

    Wei-Hong Lin, Jia-Xing Zhong, Shan Liu +2

    cs.CVcs.LGeess.IVarXiv:1911.03029v22019
  53. D-Former: A U-shaped Dilated Transformer for 3D Medical Image Segmentation

    Yixuan Wu, Kuanlun Liao, Jintai Chen +4

    cs.CVcs.AIarXiv:2201.00462v22022
  54. Three things everyone should know about Vision Transformers

    Hugo Touvron, Matthieu Cord, Alaaeldin El-Nouby +2

    cs.CVarXiv:2203.09795v12022
  55. ThinkRL-Edit: Thinking in Reinforcement Learning for Reasoning-Centric Image Editing

    Hengjia Li, Liming Jiang, Qing Yan +6

    cs.CVarXiv:2601.03467v32026
  56. Wavelet Convolutional Neural Networks

    Shin Fujieda, Kohei Takayama, Toshiya Hachisuka

    cs.CVcs.LGarXiv:1805.08620v12018
  57. LUVE : Latent-Cascaded Ultra-High-Resolution Video Generation with Dual Frequency Experts

    Chen Zhao, Jiawei Chen, Hongyu Li +6

    cs.CVarXiv:2602.11564v22026
  58. PetQA: Benchmarking Veterinary Knowledge and Clinical Reasoning

    Taegyun Kim, Youngwook Ham, Jungwook Rhim +3

    cs.CLcs.AIcs.CVarXiv:2609.04598v12026
  59. Dual-Part Multi-Lateral Branched Network for Multi-Class Segmentation in Cardiovascular Catheterization Angiograms

    Olatunji Omisore, Ahmed Elazab, Ali Shahidinejad +1

    cs.CVcs.AIcs.ROarXiv:2609.04590v12026
  60. A review on deep learning techniques for 3D sensed data classification

    David Griffiths, Jan Boehm

    cs.CVarXiv:1907.04444v12019