Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

2,401 to 2,460 of 18,817

  1. ConvMAE: Masked Convolution Meets Masked Autoencoders

    Peng Gao, Teli Ma, Hongsheng Li +3

    cs.CVarXiv:2205.03892v22022
  2. Towards Geospatial Foundation Models via Continual Pretraining

    Matias Mendieta, Boran Han, Xingjian Shi +2

    cs.CVarXiv:2302.04476v32023
  3. Overcoming Catastrophic Forgetting in Incremental Object Detection via Elastic Response Distillation

    Tao Feng, Mang Wang, Hangjie Yuan

    cs.CVarXiv:2204.02136v12022
  4. Domain Adaptive Object Detection via Asymmetric Tri-way Faster-RCNN

    Zhenwei He, Lei Zhang

    cs.CVarXiv:2007.01571v12020
  5. Improving the Performance of Unimodal Dynamic Hand-Gesture Recognition with Multimodal Training

    Mahdi Abavisani, Hamid Reza Vaezi Joze, Vishal M. Patel

    cs.CVcs.AIcs.HCarXiv:1812.06145v22018
  6. Towards Nonlinear Disentanglement in Natural Data with Temporal Sparse Coding

    David Klindt, Lukas Schott, Yash Sharma +4

    stat.MLcs.CVcs.LGarXiv:2007.10930v22020
  7. Deep Video Generation, Prediction and Completion of Human Action Sequences

    Haoye Cai, Chunyan Bai, Yu-Wing Tai +1

    cs.CVstat.MLarXiv:1711.08682v32017
  8. A Unified Continual Learning Framework with General Parameter-Efficient Tuning

    Qiankun Gao, Chen Zhao, Yifan Sun +4

    cs.CVarXiv:2303.10070v22023
  9. Cross-Domain Adaptive Clustering for Semi-Supervised Domain Adaptation

    Jichang Li, Guanbin Li, Yemin Shi +1

    cs.CVarXiv:2104.09415v12021
  10. 3D Gaussian Splatting: Survey, Technologies, Challenges, and Opportunities

    Yanqi Bao, Tianyu Ding, Jing Huo +5

    cs.CVarXiv:2407.17418v22024
  11. Gray Level Co-Occurrence Matrices: Generalisation and Some New Features

    Bino Sebastian, A. Unnikrishnan, Kannan Balakrishnan

    cs.CVarXiv:1205.4831v12012
  12. Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models

    Jiayu Wang, Yifei Ming, Zhenmei Shi +4

    cs.CVcs.AIarXiv:2406.14852v22024
  13. DeepID-Net: multi-stage and deformable deep convolutional neural networks for object detection

    Wanli Ouyang, Ping Luo, Xingyu Zeng +12

    cs.CVarXiv:1409.3505v12014
  14. Few-Shot Defect Image Generation via Defect-Aware Feature Manipulation

    Yuxuan Duan, Yan Hong, Li Niu +1

    cs.CVarXiv:2303.02389v12023
  15. Multi-Path Region Mining For Weakly Supervised 3D Semantic Segmentation on Point Clouds

    Jiacheng Wei, Guosheng Lin, Kim-Hui Yap +2

    cs.CVarXiv:2003.13035v12020
  16. TAP: Text-Aware Pre-training for Text-VQA and Text-Caption

    Zhengyuan Yang, Yijuan Lu, Jianfeng Wang +6

    cs.CVarXiv:2012.04638v12020
  17. Accel: A Corrective Fusion Network for Efficient Semantic Segmentation on Video

    Samvit Jain, Xin Wang, Joseph Gonzalez

    cs.CVcs.LGarXiv:1807.06667v42018
  18. When Does a Laugh Begin? Structured Annotator Disagreement in Temporal Laughter Localization

    Eyal Hanania, Daniel Arkushin, Naveh Ayal +4

    cs.CVcs.AIarXiv:2609.06646v12026
  19. Can we trust deep learning models diagnosis? The impact of domain shift in chest radiograph classification

    Eduardo H. P. Pooch, Pedro L. Ballester, Rodrigo C. Barros

    eess.IVcs.AIcs.CVarXiv:1909.01940v22019
  20. Diffusion Models, Image Super-Resolution And Everything: A Survey

    Brian B. Moser, Arundhati S. Shanbhag, Federico Raue +3

    cs.CVcs.AIcs.LGarXiv:2401.00736v32024
  21. Audio Surveillance: a Systematic Review

    Marco Crocco, Marco Cristani, Andrea Trucco +1

    cs.SDcs.CVcs.MMarXiv:1409.7787v12014
  22. Evaluating the Single-Shot MultiBox Detector and YOLO Deep Learning Models for the Detection of Tomatoes in a Greenhouse

    Sandro A. Magalhães, Luís Castro, Germano Moreira +4

    cs.CVcs.ROarXiv:2109.00810v12021
  23. Few-Example Object Detection with Model Communication

    Xuanyi Dong, Liang Zheng, Fan Ma +2

    cs.CVarXiv:1706.08249v82017
  24. Intra-Retinal Layer Segmentation of 3D Optical Coherence Tomography Using Coarse Grained Diffusion Map

    Raheleh Kafieh, Hossein Rabbani, Michael D. Abramoff +1

    cs.CVarXiv:1210.0310v22012
  25. Visual Explanations From Deep 3D Convolutional Neural Networks for Alzheimer's Disease Classification

    Chengliang Yang, Anand Rangarajan, Sanjay Ranka

    cs.CVcs.AIcs.LGarXiv:1803.02544v32018
  26. CoreDiff: Contextual Error-Modulated Generalized Diffusion Model for Low-Dose CT Denoising and Generalization

    Qi Gao, Zilong Li, Junping Zhang +2

    eess.IVcs.CVcs.LGarXiv:2304.01814v22023
  27. On the generalization of GAN image forensics

    Xinsheng Xuan, Bo Peng, Wei Wang +1

    cs.CVcs.LGstat.MLarXiv:1902.11153v22019
  28. Machine Vision for Natural Gas Methane Emissions Detection Using an Infrared Camera

    Jingfan Wang, Lyne P. Tchapmi, Arvind P. Ravikumara +5

    cs.CVcs.LGeess.IVarXiv:1904.08500v12019
  29. In-context learning enables multimodal large language models to classify cancer pathology images

    Dyke Ferber, Georg Wölflein, Isabella C. Wiest +8

    cs.CVarXiv:2403.07407v12024
  30. Deep Learning-Based Autonomous Driving Systems: A Survey of Attacks and Defenses

    Yao Deng, Tiehua Zhang, Guannan Lou +3

    cs.LGcs.CRcs.CVarXiv:2104.01789v22021
  31. Layer-Wise Gate-Controlled Prompt Truncation in a Multimodal Chest X-Ray Classifier

    Jingtao Lei, Hongji Li, Dexiang Shu

    cs.LGcs.AIcs.CVarXiv:2609.06590v12026
  32. Adversarial Attacks Beyond the Image Space

    Xiaohui Zeng, Chenxi Liu, Yu-Siang Wang +5

    cs.CVarXiv:1711.07183v62017
  33. Phonocardiographic Sensing using Deep Learning for Abnormal Heartbeat Detection

    Siddique Latif, Muhammad Usman, Rajib Rana +1

    cs.CVarXiv:1801.08322v42018
  34. Reading Decoder Trajectories: Training-Free Counterfactual Query-Trajectory Reliability for Small-Object Detection

    Zhaoning Shi, Bo Ma

    cs.CVcs.AIarXiv:2609.06581v12026
  35. LargeKernel3D: Scaling up Kernels in 3D Sparse CNNs

    Yukang Chen, Jianhui Liu, Xiangyu Zhang +2

    cs.CVcs.LGarXiv:2206.10555v22022
  36. OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution

    Shubhashis Roy Dipta, Sourajit Saha, Shaswati Saha +1

    cs.CVcs.AIcs.CLarXiv:2609.06490v12026
  37. Fully Convolutional One-Stage 3D Object Detection on LiDAR Range Images

    Zhi Tian, Xiangxiang Chu, Xiaoming Wang +2

    cs.CVarXiv:2205.13764v22022
  38. Branched Multi-Task Networks: Deciding What Layers To Share

    Simon Vandenhende, Stamatios Georgoulis, Bert De Brabandere +1

    cs.CVarXiv:1904.02920v52019
  39. Semi-Supervised Learning with Context-Conditional Generative Adversarial Networks

    Remi Denton, Sam Gross, Rob Fergus

    cs.CVarXiv:1611.06430v12016
  40. One MLLM, One Call: Efficient Zero-Shot Vision-and-Language Navigation via Spatial-Aware Waypoints

    Shiqi Pan, Qi Zheng, Hanqin Sun +3

    cs.CVcs.AIarXiv:2609.06476v12026
  41. Language and Visual Entity Relationship Graph for Agent Navigation

    Yicong Hong, Cristian Rodriguez-Opazo, Yuankai Qi +2

    cs.CVarXiv:2010.09304v22020
  42. Total variation regularization for fMRI-based prediction of behaviour

    Vincent Michel, Alexandre Gramfort, Gaël Varoquaux +2

    cs.CVq-bio.NCarXiv:1102.1101v12011
  43. Detail Preserved Point Cloud Completion via Separated Feature Aggregation

    Wenxiao Zhang, Qingan Yan, Chunxia Xiao

    cs.CVcs.CGarXiv:2007.02374v12020
  44. Learning Common and Specific Features for RGB-D Semantic Segmentation with Deconvolutional Networks

    Jinghua Wang, Zhenhua Wang, Dacheng Tao +2

    cs.CVarXiv:1608.01082v12016
  45. A Learned Representation for Scalable Vector Graphics

    Raphael Gontijo Lopes, David Ha, Douglas Eck +1

    cs.CVcs.LGstat.MLarXiv:1904.02632v12019
  46. Automatic Extrinsic Calibration for Lidar-Stereo Vehicle Sensor Setups

    Carlos Guindel, Jorge Beltrán, David Martín +1

    cs.CVcs.ROarXiv:1705.04085v32017
  47. Discriminative Localization in CNNs for Weakly-Supervised Segmentation of Pulmonary Nodules

    Xinyang Feng, Jie Yang, Andrew F. Laine +1

    cs.CVarXiv:1707.01086v22017
  48. Geometry Guided Adversarial Facial Expression Synthesis

    Lingxiao Song, Zhihe Lu, Ran He +2

    cs.CVarXiv:1712.03474v12017
  49. Detailed Human Shape Estimation from a Single Image by Hierarchical Mesh Deformation

    Hao Zhu, Xinxin Zuo, Sen Wang +2

    cs.CVeess.IVarXiv:1904.10506v22019
  50. The Benchmark Lottery

    Mostafa Dehghani, Yi Tay, Alexey A. Gritsenko +5

    cs.LGcs.AIcs.CLarXiv:2107.07002v12021
  51. Grounding Language Models to Images for Multimodal Inputs and Outputs

    Jing Yu Koh, Ruslan Salakhutdinov, Daniel Fried

    cs.CLcs.AIcs.CVarXiv:2301.13823v42023
  52. Adversarial Objects Against LiDAR-Based Autonomous Driving Systems

    Yulong Cao, Chaowei Xiao, Dawei Yang +4

    cs.CRcs.CVcs.LGarXiv:1907.05418v12019
  53. Spatial Information Guided Convolution for Real-Time RGBD Semantic Segmentation

    Lin-Zhuo Chen, Zheng Lin, Ziqin Wang +2

    cs.CVarXiv:2004.04534v22020
  54. Learning to Evaluate Image Captioning

    Yin Cui, Guandao Yang, Andreas Veit +2

    cs.CVcs.LGarXiv:1806.06422v12018
  55. Modeling Local Geometric Structure of 3D Point Clouds using Geo-CNN

    Shiyi Lan, Ruichi Yu, Gang Yu +1

    cs.CVarXiv:1811.07782v12018
  56. Grounding Language with Visual Affordances over Unstructured Data

    Oier Mees, Jessica Borja-Diaz, Wolfram Burgard

    cs.ROcs.AIcs.CLarXiv:2210.01911v32022
  57. Multiple Myeloma Lesion Segmentation on Whole-Body Diffusion-Weighted Imaging via Efficient Anatomical Anticipation and Multimodal Confirmation

    Mengmeng Zhang, Shengqian Huang, Junde Zhou +12

    cs.CVcs.AIarXiv:2609.06165v12026
  58. Light Field Image Super-Resolution Using Deformable Convolution

    Yingqian Wang, Jungang Yang, Longguang Wang +4

    eess.IVcs.CVarXiv:2007.03535v42020
  59. SALSA: A Novel Dataset for Multimodal Group Behavior Analysis

    Xavier Alameda-Pineda, Jacopo Staiano, Ramanathan Subramanian +5

    cs.CVarXiv:1506.06882v12015
  60. Subdivision-Based Mesh Convolution Networks

    Shi-Min Hu, Zheng-Ning Liu, Meng-Hao Guo +4

    cs.CVcs.GRcs.LGarXiv:2106.02285v22021