Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
3,541 to 3,600 of 18,822
FocalClick: Towards Practical Interactive Image Segmentation
Xi Chen, Zhiyan Zhao, Yilei Zhang +3
cs.CVarXiv:2204.02574v22022RAM: Recover Any 3D Human Motion in-the-Wild
Sen Jia, Ning Zhu, Jinqin Zhong +4
cs.CVcs.AIarXiv:2603.19929v22026The Paradigm Shift: A Comprehensive Survey on Large Vision Language Models for Multimodal Fake News Detection
Wei Ai, Yilong Tan, Yuntao Shou +4
cs.AIcs.CVarXiv:2601.15316v12026Rope3D: TheRoadside Perception Dataset for Autonomous Driving and Monocular 3D Object Detection Task
Xiaoqing Ye, Mao Shu, Hanyu Li +5
cs.CVarXiv:2203.13608v12022Visual News: Benchmark and Challenges in News Image Captioning
Fuxiao Liu, Yinghan Wang, Tianlu Wang +1
cs.CVarXiv:2010.03743v32020AdvPC: Transferable Adversarial Perturbations on 3D Point Clouds
Abdullah Hamdi, Sara Rojas, Ali Thabet +1
cs.CVcs.CRcs.LGarXiv:1912.00461v22019Image-to-Lidar Self-Supervised Distillation for Autonomous Driving Data
Corentin Sautier, Gilles Puy, Spyros Gidaris +3
cs.CVcs.LGarXiv:2203.16258v12022UNICON: Combating Label Noise Through Uniform Selection and Contrastive Learning
Nazmul Karim, Mamshad Nayeem Rizve, Nazanin Rahnavard +2
cs.CVcs.LGarXiv:2203.14542v42022Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning
Hulingxiao He, Zijun Geng, Yuxin Peng
cs.CVcs.AIarXiv:2602.07605v32026B-CNN: Branch Convolutional Neural Network for Hierarchical Classification
Xinqi Zhu, Michael Bain
cs.CVarXiv:1709.09890v22017Learning to score the figure skating sports videos
Chengming Xu, Yanwei Fu, Bing Zhang +3
cs.MMcs.CVarXiv:1802.02774v32018Nighttime Dehazing with a Synthetic Benchmark
Jing Zhang, Yang Cao, Zheng-Jun Zha +1
cs.CVcs.LGeess.IVarXiv:2008.03864v32020Rethinking Data Augmentation for Image Super-resolution: A Comprehensive Analysis and a New Strategy
Jaejun Yoo, Namhyuk Ahn, Kyung-Ah Sohn
eess.IVcs.CVarXiv:2004.00448v22020Cross-Domain Few-Shot Classification via Adversarial Task Augmentation
Haoqing Wang, Zhi-Hong Deng
cs.CVarXiv:2104.14385v22021Simple Unsupervised Object-Centric Learning for Complex and Naturalistic Videos
Gautam Singh, Yi-Fu Wu, Sungjin Ahn
cs.CVcs.LGarXiv:2205.14065v12022Learning to Hash with Binary Deep Neural Network
Thanh-Toan Do, Anh-Dzung Doan, Ngai-Man Cheung
cs.CVarXiv:1607.05140v12016MCPO: Modality-Contrastive Preference Optimization for Multimodal Chain-of-Thought Compression
Guangheng Yang, Zhenliang Ni, Zhenkai Wu +4
cs.CVcs.AIarXiv:2609.04947v12026Language-Conditioned World Modeling for Visual Navigation
Yifei Dong, Fengyi Wu, Yilong Dai +10
cs.CVcs.AIcs.ROarXiv:2603.26741v12026StableWorld: Towards Stable and Consistent Long Interactive Video Generation
Ying Yang, Zhengyao Lv, Yujia Zeng +9
cs.CVarXiv:2601.15281v22026One Diffusion Model, Two Roles: Guided Trajectory Planning and Safety-Critical Scenario Generation in Closed-Loop Simulation
Arka Pal, Rajesh Kumar, Hannes Eriksson +4
cs.CVcs.AIcs.LGarXiv:2609.04921v12026Efficient Test-Time Adaptation of Vision-Language Models
Adilbek Karmanov, Dayan Guan, Shijian Lu +2
cs.CVarXiv:2403.18293v12024Sound-based Multi-Person 3D Pose Estimation
Yusuke Oumi, Yuto Shibata, Go Irie +3
cs.CVcs.AIcs.LGarXiv:2609.04902v12026The reliability of a deep learning model in clinical out-of-distribution MRI data: a multicohort study
Gustav Mårtensson, Daniel Ferreira, Tobias Granberg +22
physics.med-phcs.CVcs.LGarXiv:1911.00515v12019Unravelling Robustness of Deep Learning based Face Recognition Against Adversarial Attacks
Gaurav Goswami, Nalini Ratha, Akshay Agarwal +2
cs.CVarXiv:1803.00401v12018Highly Accurate Dichotomous Image Segmentation
Xuebin Qin, Hang Dai, Xiaobin Hu +3
cs.CVarXiv:2203.03041v42022Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form Sentences
Zhu Zhang, Zhou Zhao, Yang Zhao +3
cs.CVarXiv:2001.06891v32020LVSM: A Large View Synthesis Model with Minimal 3D Inductive Bias
Haian Jin, Hanwen Jiang, Hao Tan +6
cs.CVcs.GRcs.LGarXiv:2410.17242v22024Parametric Classification for Generalized Category Discovery: A Baseline Study
Xin Wen, Bingchen Zhao, Xiaojuan Qi
cs.CVcs.LGarXiv:2211.11727v42022Methane Detection On Board Satellites from Unorthorectified Imagery
Luca Marini, Maggie Chen, Hala Lamdouar +3
cs.CVcs.AIcs.LGarXiv:2609.04906v12026A very preliminary analysis of DALL-E 2
Gary Marcus, Ernest Davis, Scott Aaronson
cs.CVcs.AIarXiv:2204.13807v22022SimFuse3D: Source-Guided Target Simulation and Confidence-Guided Multi-Stage Localization Reweighting for Cross-Platform 3D Object Detection
Yongchun Lin, Xinliang Zhang, Yun Zou +7
cs.CVcs.AIarXiv:2609.04886v12026Foveation-based Mechanisms Alleviate Adversarial Examples
Yan Luo, Xavier Boix, Gemma Roig +2
cs.LGcs.CVarXiv:1511.06292v32015Mitigating Performance Discrepancy in Cross-Domain 3D Class-Incremental Learning
Jinge Ma, Gautham Vinod, Bruce Coburn +3
cs.CVcs.AIarXiv:2609.04860v12026Adversarial Example Detection for DNN Models: A Review and Experimental Comparison
Ahmed Aldahdooh, Wassim Hamidouche, Sid Ahmed Fezza +1
cs.CVcs.CRarXiv:2105.00203v42021VPFNet: Improving 3D Object Detection with Virtual Point based LiDAR and Stereo Data Fusion
Hanqi Zhu, Jiajun Deng, Yu Zhang +4
cs.CVarXiv:2111.14382v22021Not All Directions Matter: Towards Structured and Task-Aware Low-Rank Model Adaptation
Xi Xiao, Chenrui Ma, Yunbei Zhang +7
cs.CVarXiv:2603.14228v22026SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
Jian Zhang, Shijie Zhou, Bangya Liu +2
cs.CVarXiv:2603.27437v32026DealMVC: Dual Contrastive Calibration for Multi-view Clustering
Xihong Yang, Jiaqi Jin, Siwei Wang +7
cs.CVcs.LGarXiv:2308.09000v32023Linguistic Trajectory Encoding for Efficient Long-Horizon Spatial Memory in Embodied Agents
Tianyidan Xie, Shenyi Wang, Qiang Tang +7
cs.CVcs.AIarXiv:2609.04802v12026Satellite Imagery Feature Detection using Deep Convolutional Neural Network: A Kaggle Competition
Vladimir Iglovikov, Sergey Mushinskiy, Vladimir Osin
cs.CVarXiv:1706.06169v12017Pansharpening for Thin-Cloud Contaminated Remote Sensing Images: A Unified Framework and Benchmark Dataset
Songcheng Du, Yang Zou, Jiaxin Li +4
cs.CVarXiv:2603.14952v12026Dynamic High-frequency Convolution for Infrared Small Target Detection
Ruojing Li, Chao Xiao, Qian Yin +5
cs.CVarXiv:2602.02969v22026Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild
Changda Zhou, Ziyue Gao, Xueqing Wang +4
cs.CVarXiv:2603.04205v22026HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis
Mingjin Chen, Junhao Chen, Zhaoxin Fan +6
cs.CVarXiv:2604.03305v12026MR Image Denoising and Super-Resolution Using Regularized Reverse Diffusion
Hyungjin Chung, Eun Sun Lee, Jong Chul Ye
eess.IVcs.AIcs.CVarXiv:2203.12621v12022Efficient Medical Image Segmentation Based on Knowledge Distillation
Dian Qin, Jiajun Bu, Zhe Liu +6
eess.IVcs.CVarXiv:2108.09987v12021Unsupervised Domain Adaptation using Feature-Whitening and Consensus Loss
Subhankar Roy, Aliaksandr Siarohin, Enver Sangineto +3
cs.CVarXiv:1903.03215v22019DAE-Former: Dual Attention-guided Efficient Transformer for Medical Image Segmentation
Reza Azad, René Arimond, Ehsan Khodapanah Aghdam +2
cs.CVarXiv:2212.13504v32022SAM-CLIP: Merging Vision Foundation Models towards Semantic and Spatial Understanding
Haoxiang Wang, Pavan Kumar Anasosalu Vasu, Fartash Faghri +6
cs.CVcs.LGarXiv:2310.15308v42023Convolutional Kolmogorov-Arnold Networks
Alexander Dylan Bodner, Antonio Santiago Tepsich, Jack Natan Spolski +1
cs.CVcs.AIarXiv:2406.13155v32024Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies
Wenjin Hou, Wei Liu, Han Hu +3
cs.CVarXiv:2602.01816v12026RoIMix: Proposal-Fusion among Multiple Images for Underwater Object Detection
Wei-Hong Lin, Jia-Xing Zhong, Shan Liu +2
cs.CVcs.LGeess.IVarXiv:1911.03029v22019D-Former: A U-shaped Dilated Transformer for 3D Medical Image Segmentation
Yixuan Wu, Kuanlun Liao, Jintai Chen +4
cs.CVcs.AIarXiv:2201.00462v22022Three things everyone should know about Vision Transformers
Hugo Touvron, Matthieu Cord, Alaaeldin El-Nouby +2
cs.CVarXiv:2203.09795v12022ThinkRL-Edit: Thinking in Reinforcement Learning for Reasoning-Centric Image Editing
Hengjia Li, Liming Jiang, Qing Yan +6
cs.CVarXiv:2601.03467v32026Wavelet Convolutional Neural Networks
Shin Fujieda, Kohei Takayama, Toshiya Hachisuka
cs.CVcs.LGarXiv:1805.08620v12018LUVE : Latent-Cascaded Ultra-High-Resolution Video Generation with Dual Frequency Experts
Chen Zhao, Jiawei Chen, Hongyu Li +6
cs.CVarXiv:2602.11564v22026PetQA: Benchmarking Veterinary Knowledge and Clinical Reasoning
Taegyun Kim, Youngwook Ham, Jungwook Rhim +3
cs.CLcs.AIcs.CVarXiv:2609.04598v12026Dual-Part Multi-Lateral Branched Network for Multi-Class Segmentation in Cardiovascular Catheterization Angiograms
Olatunji Omisore, Ahmed Elazab, Ali Shahidinejad +1
cs.CVcs.AIcs.ROarXiv:2609.04590v12026A review on deep learning techniques for 3D sensed data classification
David Griffiths, Jan Boehm
cs.CVarXiv:1907.04444v12019