Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

2,341 to 2,400 of 18,786

  1. SegSort: Segmentation by Discriminative Sorting of Segments

    Jyh-Jing Hwang, Stella X. Yu, Jianbo Shi +4

    cs.CVcs.LGeess.IVarXiv:1910.06962v22019
  2. XIRL: Cross-embodiment Inverse Reinforcement Learning

    Kevin Zakka, Andy Zeng, Pete Florence +3

    cs.ROcs.AIcs.CVarXiv:2106.03911v32021
  3. SparseFusion: Fusing Multi-Modal Sparse Representations for Multi-Sensor 3D Object Detection

    Yichen Xie, Chenfeng Xu, Marie-Julie Rakotosaona +5

    cs.CVarXiv:2304.14340v12023
  4. Comparative Study of Anatomical and Learned Features in AI Models for Structural Brain MRI

    Boyang Yu, Miquel Lopez Escoriza, Long Chen +3

    cs.CVcs.AIarXiv:2609.06807v12026
  5. CompGS: Smaller and Faster Gaussian Splatting with Vector Quantization

    KL Navaneet, Kossar Pourahmadi Meibodi, Soroush Abbasi Koohpayegani +1

    cs.CVarXiv:2311.18159v32023
  6. Blended Multi-Modal Deep ConvNet Features for Diabetic Retinopathy Severity Prediction

    J. D. Bodapati, N. Veeranjaneyulu, S. N. Shareef +4

    eess.IVcs.CVcs.LGarXiv:2006.00197v12020
  7. Structured Binary Neural Networks for Accurate Image Classification and Semantic Segmentation

    Bohan Zhuang, Chunhua Shen, Mingkui Tan +2

    cs.CVarXiv:1811.10413v22018
  8. Mumford-Shah Loss Functional for Image Segmentation with Deep Learning

    Boah Kim, Jong Chul Ye

    cs.CVcs.LGstat.MLarXiv:1904.02872v22019
  9. Learning to Reconstruct Shapes from Unseen Classes

    Xiuming Zhang, Zhoutong Zhang, Chengkai Zhang +3

    cs.CVcs.AIarXiv:1812.11166v12018
  10. When 3D Gaussian Splatting Recovers Real Surfaces

    Songhe Wang, David Johnathan Miller

    cs.LGcs.CVarXiv:2608.30054v12026
  11. Boundary Voting Network for Ambiguity-Aware Timestamp-Supervised Action Segmentation

    Runzhong Zhang, Yueqi Duan, Yang Chen +4

    cs.CVarXiv:2609.08167v12026
  12. se(3)-TrackNet: Data-driven 6D Pose Tracking by Calibrating Image Residuals in Synthetic Domains

    Bowen Wen, Chaitanya Mitash, Baozhang Ren +1

    cs.CVcs.GRcs.LGarXiv:2007.13866v12020
  13. Unified Quality Assessment of In-the-Wild Videos with Mixed Datasets Training

    Dingquan Li, Tingting Jiang, Ming Jiang

    cs.CVcs.MMeess.IVarXiv:2011.04263v22020
  14. Selective Refinement Network for High Performance Face Detection

    Cheng Chi, Shifeng Zhang, Junliang Xing +3

    cs.CVarXiv:1809.02693v12018
  15. Identity-Preserving Text-to-Video Generation by Frequency Decomposition

    Shenghai Yuan, Jinfa Huang, Xianyi He +5

    cs.CVcs.MMarXiv:2411.17440v32024
  16. Pyramid Diffusion Models For Low-light Image Enhancement

    Dewei Zhou, Zongxin Yang, Yi Yang

    cs.CVarXiv:2305.10028v12023
  17. FSAN: Flow State Attention Network for Aerodynamic Prediction

    Wenxuan Jin, Jianguo Yao, Haibing Guan +1

    cs.CVcs.AIarXiv:2609.06660v12026
  18. ECOKV: Geometry-Aware KV Cache Eviction via Complementary Diversity Metrics

    Chin Ting Hsu, Yu-Syuan Xu, Ling Zou +2

    cs.CLcs.AIcs.CVarXiv:2609.06663v12026
  19. Deep Regression Forests for Age Estimation

    Wei Shen, Yilu Guo, Yan Wang +3

    cs.CVarXiv:1712.07195v12017
  20. Debiased Learning from Naturally Imbalanced Pseudo-Labels

    Xudong Wang, Zhirong Wu, Long Lian +1

    cs.LGcs.CLcs.CVarXiv:2201.01490v22022
  21. Robust Training under Label Noise by Over-parameterization

    Sheng Liu, Zhihui Zhu, Qing Qu +1

    cs.LGcs.AIcs.CVarXiv:2202.14026v22022
  22. Unsupervised Meta-Learning For Few-Shot Image Classification

    Siavash Khodadadeh, Ladislau Bölöni, Mubarak Shah

    cs.CVcs.LGarXiv:1811.11819v22018
  23. Deep Sinogram Completion with Image Prior for Metal Artifact Reduction in CT Images

    Lequan Yu, Zhicheng Zhang, Xiaomeng Li +1

    eess.IVcs.CVarXiv:2009.07469v12020
  24. Dictionary Learning and Sparse Coding on Grassmann Manifolds: An Extrinsic Solution

    Mehrtash Harandi, Conrad Sanderson, Chunhua Shen +1

    cs.CVarXiv:1310.4891v12013
  25. Semantic Jitter: Dense Supervision for Visual Comparisons via Synthetic Images

    Aron Yu, Kristen Grauman

    cs.CVarXiv:1612.06341v22016
  26. Companion-style QA Assistance in Ego-Vision

    Hangyu Qin, Junbin Xiao, Shenglang Zhang +1

    cs.CVcs.AIarXiv:2609.06721v12026
  27. Counterfactual Tests for Measuring Chain-of-Thought Faithfulness in Visual Language Models

    Bayar Menzat, Maximilian Süss, Ruizhi Wang +3

    cs.CVcs.AIcs.CLarXiv:2609.06704v12026
  28. FashionBERT: Text and Image Matching with Adaptive Loss for Cross-modal Retrieval

    Dehong Gao, Linbo Jin, Ben Chen +5

    cs.IRcs.CVcs.LGarXiv:2005.09801v22020
  29. Attention-Enhanced Deep Features with Heterogeneous Ensemble Learning for Glaucoma Detection

    Abdullah Al Shafi, Nishat Sadaf Lira, Abrar Hasan +2

    cs.CVcs.AIcs.LGarXiv:2609.06699v12026
  30. ConvMAE: Masked Convolution Meets Masked Autoencoders

    Peng Gao, Teli Ma, Hongsheng Li +3

    cs.CVarXiv:2205.03892v22022
  31. Towards Geospatial Foundation Models via Continual Pretraining

    Matias Mendieta, Boran Han, Xingjian Shi +2

    cs.CVarXiv:2302.04476v32023
  32. Overcoming Catastrophic Forgetting in Incremental Object Detection via Elastic Response Distillation

    Tao Feng, Mang Wang, Hangjie Yuan

    cs.CVarXiv:2204.02136v12022
  33. Domain Adaptive Object Detection via Asymmetric Tri-way Faster-RCNN

    Zhenwei He, Lei Zhang

    cs.CVarXiv:2007.01571v12020
  34. Improving the Performance of Unimodal Dynamic Hand-Gesture Recognition with Multimodal Training

    Mahdi Abavisani, Hamid Reza Vaezi Joze, Vishal M. Patel

    cs.CVcs.AIcs.HCarXiv:1812.06145v22018
  35. Towards Nonlinear Disentanglement in Natural Data with Temporal Sparse Coding

    David Klindt, Lukas Schott, Yash Sharma +4

    stat.MLcs.CVcs.LGarXiv:2007.10930v22020
  36. Deep Video Generation, Prediction and Completion of Human Action Sequences

    Haoye Cai, Chunyan Bai, Yu-Wing Tai +1

    cs.CVstat.MLarXiv:1711.08682v32017
  37. A Unified Continual Learning Framework with General Parameter-Efficient Tuning

    Qiankun Gao, Chen Zhao, Yifan Sun +4

    cs.CVarXiv:2303.10070v22023
  38. Cross-Domain Adaptive Clustering for Semi-Supervised Domain Adaptation

    Jichang Li, Guanbin Li, Yemin Shi +1

    cs.CVarXiv:2104.09415v12021
  39. 3D Gaussian Splatting: Survey, Technologies, Challenges, and Opportunities

    Yanqi Bao, Tianyu Ding, Jing Huo +5

    cs.CVarXiv:2407.17418v22024
  40. Gray Level Co-Occurrence Matrices: Generalisation and Some New Features

    Bino Sebastian, A. Unnikrishnan, Kannan Balakrishnan

    cs.CVarXiv:1205.4831v12012
  41. Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models

    Jiayu Wang, Yifei Ming, Zhenmei Shi +4

    cs.CVcs.AIarXiv:2406.14852v22024
  42. DeepID-Net: multi-stage and deformable deep convolutional neural networks for object detection

    Wanli Ouyang, Ping Luo, Xingyu Zeng +12

    cs.CVarXiv:1409.3505v12014
  43. Few-Shot Defect Image Generation via Defect-Aware Feature Manipulation

    Yuxuan Duan, Yan Hong, Li Niu +1

    cs.CVarXiv:2303.02389v12023
  44. Multi-Path Region Mining For Weakly Supervised 3D Semantic Segmentation on Point Clouds

    Jiacheng Wei, Guosheng Lin, Kim-Hui Yap +2

    cs.CVarXiv:2003.13035v12020
  45. TAP: Text-Aware Pre-training for Text-VQA and Text-Caption

    Zhengyuan Yang, Yijuan Lu, Jianfeng Wang +6

    cs.CVarXiv:2012.04638v12020
  46. Accel: A Corrective Fusion Network for Efficient Semantic Segmentation on Video

    Samvit Jain, Xin Wang, Joseph Gonzalez

    cs.CVcs.LGarXiv:1807.06667v42018
  47. When Does a Laugh Begin? Structured Annotator Disagreement in Temporal Laughter Localization

    Eyal Hanania, Daniel Arkushin, Naveh Ayal +4

    cs.CVcs.AIarXiv:2609.06646v12026
  48. Can we trust deep learning models diagnosis? The impact of domain shift in chest radiograph classification

    Eduardo H. P. Pooch, Pedro L. Ballester, Rodrigo C. Barros

    eess.IVcs.AIcs.CVarXiv:1909.01940v22019
  49. Diffusion Models, Image Super-Resolution And Everything: A Survey

    Brian B. Moser, Arundhati S. Shanbhag, Federico Raue +3

    cs.CVcs.AIcs.LGarXiv:2401.00736v32024
  50. Audio Surveillance: a Systematic Review

    Marco Crocco, Marco Cristani, Andrea Trucco +1

    cs.SDcs.CVcs.MMarXiv:1409.7787v12014
  51. Evaluating the Single-Shot MultiBox Detector and YOLO Deep Learning Models for the Detection of Tomatoes in a Greenhouse

    Sandro A. Magalhães, Luís Castro, Germano Moreira +4

    cs.CVcs.ROarXiv:2109.00810v12021
  52. Few-Example Object Detection with Model Communication

    Xuanyi Dong, Liang Zheng, Fan Ma +2

    cs.CVarXiv:1706.08249v82017
  53. Intra-Retinal Layer Segmentation of 3D Optical Coherence Tomography Using Coarse Grained Diffusion Map

    Raheleh Kafieh, Hossein Rabbani, Michael D. Abramoff +1

    cs.CVarXiv:1210.0310v22012
  54. Visual Explanations From Deep 3D Convolutional Neural Networks for Alzheimer's Disease Classification

    Chengliang Yang, Anand Rangarajan, Sanjay Ranka

    cs.CVcs.AIcs.LGarXiv:1803.02544v32018
  55. CoreDiff: Contextual Error-Modulated Generalized Diffusion Model for Low-Dose CT Denoising and Generalization

    Qi Gao, Zilong Li, Junping Zhang +2

    eess.IVcs.CVcs.LGarXiv:2304.01814v22023
  56. On the generalization of GAN image forensics

    Xinsheng Xuan, Bo Peng, Wei Wang +1

    cs.CVcs.LGstat.MLarXiv:1902.11153v22019
  57. Machine Vision for Natural Gas Methane Emissions Detection Using an Infrared Camera

    Jingfan Wang, Lyne P. Tchapmi, Arvind P. Ravikumara +5

    cs.CVcs.LGeess.IVarXiv:1904.08500v12019
  58. In-context learning enables multimodal large language models to classify cancer pathology images

    Dyke Ferber, Georg Wölflein, Isabella C. Wiest +8

    cs.CVarXiv:2403.07407v12024
  59. Deep Learning-Based Autonomous Driving Systems: A Survey of Attacks and Defenses

    Yao Deng, Tiehua Zhang, Guannan Lou +3

    cs.LGcs.CRcs.CVarXiv:2104.01789v22021
  60. Layer-Wise Gate-Controlled Prompt Truncation in a Multimodal Chest X-Ray Classifier

    Jingtao Lei, Hongji Li, Dexiang Shu

    cs.LGcs.AIcs.CVarXiv:2609.06590v12026