Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

14,041 to 14,100 of 18,830

  1. NetAdapt: Platform-Aware Neural Network Adaptation for Mobile Applications

    Tien-Ju Yang, Andrew Howard, Bo Chen +5

    cs.CVarXiv:1804.03230v22018
  2. PatchGate: Narrowing the Verbalization Gap with Intrinsic Object Inventories in Frozen Vision-Language Models

    Jihyung Ko, Eunji Jung, Hyeongsub Kim +4

    cs.CVcs.AIcs.CLarXiv:2608.21819v12026
  3. LiteEvent-AE: Lightweight Autoencoder for Event-Based Vision on Low-Latency Energy-Constrained Edge Devices

    Riadul Islam, Joey Mule, Dhandeep Challagundla +3

    cs.CVcs.AIeess.IVarXiv:2608.21764v12026
  4. Single-Image Depth Perception in the Wild

    Weifeng Chen, Zhao Fu, Dawei Yang +1

    cs.CVcs.AIarXiv:1604.03901v22016
  5. Searching Central Difference Convolutional Networks for Face Anti-Spoofing

    Zitong Yu, Chenxu Zhao, Zezheng Wang +5

    cs.CVarXiv:2003.04092v12020
  6. GuidedFlow: An Attention-Guided Framework for Anomaly Detection in Additive Manufacturing

    Sosmita Paul, Krishna Roy

    cs.CVcs.LGarXiv:2608.22789v12026
  7. Material Recognition in the Wild with the Materials in Context Database

    Sean Bell, Paul Upchurch, Noah Snavely +1

    cs.CVarXiv:1412.0623v22014
  8. MDFI: A Multi-Domain Features Integration for Compressed Video Quality Enhancement

    Sang NguyenQuang, Hieu Bui Minh, Dang BuiDinh +1

    eess.IVcs.CVarXiv:2608.21495v12026
  9. What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model Analysis

    Jeonghun Baek, Geewook Kim, Junyeop Lee +5

    cs.CVarXiv:1904.01906v42019
  10. A Simulator-Grounded Framework For Constructing Verifiable Muscle-Grounded QA From 3D Tongue Meshes (extended version)

    Seungho Eum, Unsang Park

    cs.CVcs.HCarXiv:2608.23137v22026
  11. YouTube-BoundingBoxes: A Large High-Precision Human-Annotated Data Set for Object Detection in Video

    Esteban Real, Jonathon Shlens, Stefano Mazzocchi +2

    cs.CVarXiv:1702.00824v52017
  12. Data-free parameter pruning for Deep Neural Networks

    Suraj Srinivas, R. Venkatesh Babu

    cs.CVarXiv:1507.06149v12015
  13. Automatic Knee Osteoarthritis Diagnosis from Plain Radiographs: A Deep Learning-Based Approach

    Aleksei Tiulpin, Jérôme Thevenot, Esa Rahtu +2

    cs.CVarXiv:1710.10589v12017
  14. Robust Image Sentiment Analysis Using Progressively Trained and Domain Transferred Deep Networks

    Quanzeng You, Jiebo Luo, Hailin Jin +1

    cs.CVcs.IRcs.LGarXiv:1509.06041v12015
  15. FlashReg: GPU-Accelerated 3-Clique Point Cloud Registration for Real-Time Correspondence-to-Pose Estimation

    Ziyang Yu, Xiang Li, Qiong Chang +1

    cs.CVcs.DCarXiv:2608.21804v12026
  16. Frame-Level Evaluation in Weakly Supervised Video Anomaly Detection Mostly Measures Video-Level Ranking

    Inpyo Song, Jangwon Lee

    cs.CVarXiv:2608.21854v12026
  17. Revisiting Dilated Convolution: A Simple Approach for Weakly- and Semi- Supervised Semantic Segmentation

    Yunchao Wei, Huaxin Xiao, Honghui Shi +3

    cs.CVarXiv:1805.04574v22018
  18. GaussVid: Sparse-View Gaussian Splatting with 3D-Aware Video Diffusion Priors

    Xinhui Liu, Can Wang, Wei Jiang +2

    cs.CVarXiv:2608.21849v12026
  19. V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning

    Shulin Tian, Minglun Li, Yuhao Dong +6

    cs.CVcs.AIarXiv:2608.25580v12026
  20. Depth-Aware Video Frame Interpolation

    Wenbo Bao, Wei-Sheng Lai, Chao Ma +3

    cs.CVarXiv:1904.00830v12019
  21. Towards Generalist Biomedical AI

    Tao Tu, Shekoofeh Azizi, Danny Driess +29

    cs.CLcs.CVarXiv:2307.14334v12023
  22. Unsupervised Learning of Disentangled Representations from Video

    Remi Denton, Vighnesh Birodkar

    cs.LGcs.AIcs.CVarXiv:1705.10915v12017
  23. CompressAI: a PyTorch library and evaluation platform for end-to-end compression research

    Jean Bégaint, Fabien Racapé, Simon Feltman +1

    cs.CVeess.IVarXiv:2011.03029v12020
  24. Framing U-Net via Deep Convolutional Framelets: Application to Sparse-view CT

    Yoseob Han, Jong Chul Ye

    cs.CVcs.LGstat.MLarXiv:1708.08333v32017
  25. AudioCLIP: Extending CLIP to Image, Text and Audio

    Andrey Guzhov, Federico Raue, Jörn Hees +1

    cs.SDcs.CVeess.ASarXiv:2106.13043v12021
  26. Video Paragraph Captioning Using Hierarchical Recurrent Neural Networks

    Haonan Yu, Jiang Wang, Zhiheng Huang +2

    cs.CVarXiv:1510.07712v22015
  27. Towards Optimal Structured CNN Pruning via Generative Adversarial Learning

    Shaohui Lin, Rongrong Ji, Chenqian Yan +5

    cs.CVarXiv:1903.09291v12019
  28. DefaultShift: Auditing Semantic Default Shift in Accelerated Text-to-Image Models

    Xuanhua Yin, Chuanzhi Xu, Shunqi Mao +2

    cs.CVarXiv:2608.21784v12026
  29. Improving Computer-aided Detection using Convolutional Neural Networks and Random View Aggregation

    Holger R. Roth, Le Lu, Jiamin Liu +5

    cs.CVarXiv:1505.03046v22015
  30. Calibrate What You SHIP: Post-Selection Risk Control for Verifier-Guided Text-to-Image Generation

    Xuanhua Yin, Shunqi Mao, Wei Guo +2

    cs.CVarXiv:2608.21748v12026
  31. Geometric GAN

    Jae Hyun Lim, Jong Chul Ye

    stat.MLcond-mat.dis-nncs.AIarXiv:1705.02894v22017
  32. Learning to Compose Neural Networks for Question Answering

    Jacob Andreas, Marcus Rohrbach, Trevor Darrell +1

    cs.CLcs.CVcs.NEarXiv:1601.01705v42016
  33. Attend, Infer, Repeat: Fast Scene Understanding with Generative Models

    S. M. Ali Eslami, Nicolas Heess, Theophane Weber +4

    cs.CVcs.LGarXiv:1603.08575v32016
  34. f-VAEGAN-D2: A Feature Generating Framework for Any-Shot Learning

    Yongqin Xian, Saurabh Sharma, Bernt Schiele +1

    cs.CVarXiv:1903.10132v12019
  35. Phenaki: Variable Length Video Generation From Open Domain Textual Description

    Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans +6

    cs.CVcs.AIarXiv:2210.02399v12022
  36. A Multi-View Embedding Space for Modeling Internet Images, Tags, and their Semantics

    Yunchao Gong, Qifa Ke, Michael Isard +1

    cs.CVcs.IRcs.LGarXiv:1212.4522v22012
  37. TextCaps: a Dataset for Image Captioning with Reading Comprehension

    Oleksii Sidorov, Ronghang Hu, Marcus Rohrbach +1

    cs.CVcs.CLarXiv:2003.12462v22020
  38. Flowing ConvNets for Human Pose Estimation in Videos

    Tomas Pfister, James Charles, Andrew Zisserman

    cs.CVarXiv:1506.02897v22015
  39. Co-occurrence Feature Learning from Skeleton Data for Action Recognition and Detection with Hierarchical Aggregation

    Chao Li, Qiaoyong Zhong, Di Xie +1

    cs.CVarXiv:1804.06055v12018
  40. V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

    Penghao Wu, Saining Xie

    cs.CVarXiv:2312.14135v22023
  41. Unite the People: Closing the Loop Between 3D and 2D Human Representations

    Christoph Lassner, Javier Romero, Martin Kiefel +3

    cs.CVarXiv:1701.02468v32017
  42. Convolutional neural network architecture for geometric matching

    Ignacio Rocco, Relja Arandjelović, Josef Sivic

    cs.CVcs.LGarXiv:1703.05593v22017
  43. Inferring and Executing Programs for Visual Reasoning

    Justin Johnson, Bharath Hariharan, Laurens van der Maaten +4

    cs.CVcs.CLcs.LGarXiv:1705.03633v12017
  44. Pretreatment DCE-MRI Resolves Response Quality Within Pathologic Endpoints in Neoadjuvant Breast Cancer

    Dattatreya Kantha, Murray H. Loew

    eess.IVcs.CVcs.LGarXiv:2608.22097v12026
  45. Input-Aware Dynamic Backdoor Attack

    Anh Nguyen, Anh Tran

    cs.CRcs.CVarXiv:2010.08138v12020
  46. End-to-end people detection in crowded scenes

    Russell Stewart, Mykhaylo Andriluka

    cs.CVarXiv:1506.04878v32015
  47. DIRE for Diffusion-Generated Image Detection

    Zhendong Wang, Jianmin Bao, Wengang Zhou +4

    cs.CVarXiv:2303.09295v12023
  48. Binary Neural Networks: A Survey

    Haotong Qin, Ruihao Gong, Xianglong Liu +3

    cs.NEcs.CVcs.LGarXiv:2004.03333v12020
  49. VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

    Haodong Duan, Xinyu Fang, Junming Yang +41

    cs.CVarXiv:2407.11691v52024
  50. A New 2.5D Representation for Lymph Node Detection using Random Sets of Deep Convolutional Neural Network Observations

    Holger R. Roth, Le Lu, Ari Seff +6

    cs.CVcs.LGcs.NEarXiv:1406.2639v12014
  51. Learning a Multi-View Stereo Machine

    Abhishek Kar, Christian Häne, Jitendra Malik

    cs.CVarXiv:1708.05375v12017
  52. Spiking Neural Networks for Energy-Efficient Object Detection in Forward-Looking Sonar Imagery

    Gwenevere Frank, Gert Cauwenberghs

    cs.CVeess.SParXiv:2608.22072v12026
  53. Low-Light Image and Video Enhancement Using Deep Learning: A Survey

    Chongyi Li, Chunle Guo, Linghao Han +4

    cs.CVarXiv:2104.10729v32021
  54. VIG: Visual Information Gain as a Reward Signal for Multimodal Chain-of-Thought Compression

    Wen Luo, Xiaohan Yi, Xiaotao Huang +1

    cs.CVarXiv:2608.21883v12026
  55. Stochastic Variational Video Prediction

    Mohammad Babaeizadeh, Chelsea Finn, Dumitru Erhan +2

    cs.CVcs.ROarXiv:1710.11252v22017
  56. Image Captioning: Transforming Objects into Words

    Simao Herdade, Armin Kappeler, Kofi Boakye +1

    cs.CVcs.CLarXiv:1906.05963v22019
  57. Visual Saliency Based on Scale-Space Analysis in the Frequency Domain

    Jian Li, Martin Levine, Xiangjing An +2

    cs.CVarXiv:1605.01999v12016
  58. A Survey of Recent Advances in CNN-based Single Image Crowd Counting and Density Estimation

    Vishwanath A. Sindagi, Vishal M. Patel

    cs.CVarXiv:1707.01202v12017
  59. Learning Human-Object Interactions by Graph Parsing Neural Networks

    Siyuan Qi, Wenguan Wang, Baoxiong Jia +2

    cs.CVarXiv:1808.07962v12018
  60. Pointwise Convolutional Neural Networks

    Binh-Son Hua, Minh-Khoi Tran, Sai-Kit Yeung

    cs.CVcs.LGarXiv:1712.05245v22017