Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
14,041 to 14,100 of 18,830
NetAdapt: Platform-Aware Neural Network Adaptation for Mobile Applications
Tien-Ju Yang, Andrew Howard, Bo Chen +5
cs.CVarXiv:1804.03230v22018PatchGate: Narrowing the Verbalization Gap with Intrinsic Object Inventories in Frozen Vision-Language Models
Jihyung Ko, Eunji Jung, Hyeongsub Kim +4
cs.CVcs.AIcs.CLarXiv:2608.21819v12026LiteEvent-AE: Lightweight Autoencoder for Event-Based Vision on Low-Latency Energy-Constrained Edge Devices
Riadul Islam, Joey Mule, Dhandeep Challagundla +3
cs.CVcs.AIeess.IVarXiv:2608.21764v12026Single-Image Depth Perception in the Wild
Weifeng Chen, Zhao Fu, Dawei Yang +1
cs.CVcs.AIarXiv:1604.03901v22016Searching Central Difference Convolutional Networks for Face Anti-Spoofing
Zitong Yu, Chenxu Zhao, Zezheng Wang +5
cs.CVarXiv:2003.04092v12020GuidedFlow: An Attention-Guided Framework for Anomaly Detection in Additive Manufacturing
Sosmita Paul, Krishna Roy
cs.CVcs.LGarXiv:2608.22789v12026Material Recognition in the Wild with the Materials in Context Database
Sean Bell, Paul Upchurch, Noah Snavely +1
cs.CVarXiv:1412.0623v22014MDFI: A Multi-Domain Features Integration for Compressed Video Quality Enhancement
Sang NguyenQuang, Hieu Bui Minh, Dang BuiDinh +1
eess.IVcs.CVarXiv:2608.21495v12026What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model Analysis
Jeonghun Baek, Geewook Kim, Junyeop Lee +5
cs.CVarXiv:1904.01906v42019A Simulator-Grounded Framework For Constructing Verifiable Muscle-Grounded QA From 3D Tongue Meshes (extended version)
Seungho Eum, Unsang Park
cs.CVcs.HCarXiv:2608.23137v22026YouTube-BoundingBoxes: A Large High-Precision Human-Annotated Data Set for Object Detection in Video
Esteban Real, Jonathon Shlens, Stefano Mazzocchi +2
cs.CVarXiv:1702.00824v52017Data-free parameter pruning for Deep Neural Networks
Suraj Srinivas, R. Venkatesh Babu
cs.CVarXiv:1507.06149v12015Automatic Knee Osteoarthritis Diagnosis from Plain Radiographs: A Deep Learning-Based Approach
Aleksei Tiulpin, Jérôme Thevenot, Esa Rahtu +2
cs.CVarXiv:1710.10589v12017Robust Image Sentiment Analysis Using Progressively Trained and Domain Transferred Deep Networks
Quanzeng You, Jiebo Luo, Hailin Jin +1
cs.CVcs.IRcs.LGarXiv:1509.06041v12015FlashReg: GPU-Accelerated 3-Clique Point Cloud Registration for Real-Time Correspondence-to-Pose Estimation
Ziyang Yu, Xiang Li, Qiong Chang +1
cs.CVcs.DCarXiv:2608.21804v12026Frame-Level Evaluation in Weakly Supervised Video Anomaly Detection Mostly Measures Video-Level Ranking
Inpyo Song, Jangwon Lee
cs.CVarXiv:2608.21854v12026Revisiting Dilated Convolution: A Simple Approach for Weakly- and Semi- Supervised Semantic Segmentation
Yunchao Wei, Huaxin Xiao, Honghui Shi +3
cs.CVarXiv:1805.04574v22018GaussVid: Sparse-View Gaussian Splatting with 3D-Aware Video Diffusion Priors
Xinhui Liu, Can Wang, Wei Jiang +2
cs.CVarXiv:2608.21849v12026V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning
Shulin Tian, Minglun Li, Yuhao Dong +6
cs.CVcs.AIarXiv:2608.25580v12026Depth-Aware Video Frame Interpolation
Wenbo Bao, Wei-Sheng Lai, Chao Ma +3
cs.CVarXiv:1904.00830v12019Towards Generalist Biomedical AI
Tao Tu, Shekoofeh Azizi, Danny Driess +29
cs.CLcs.CVarXiv:2307.14334v12023Unsupervised Learning of Disentangled Representations from Video
Remi Denton, Vighnesh Birodkar
cs.LGcs.AIcs.CVarXiv:1705.10915v12017CompressAI: a PyTorch library and evaluation platform for end-to-end compression research
Jean Bégaint, Fabien Racapé, Simon Feltman +1
cs.CVeess.IVarXiv:2011.03029v12020Framing U-Net via Deep Convolutional Framelets: Application to Sparse-view CT
Yoseob Han, Jong Chul Ye
cs.CVcs.LGstat.MLarXiv:1708.08333v32017AudioCLIP: Extending CLIP to Image, Text and Audio
Andrey Guzhov, Federico Raue, Jörn Hees +1
cs.SDcs.CVeess.ASarXiv:2106.13043v12021Video Paragraph Captioning Using Hierarchical Recurrent Neural Networks
Haonan Yu, Jiang Wang, Zhiheng Huang +2
cs.CVarXiv:1510.07712v22015Towards Optimal Structured CNN Pruning via Generative Adversarial Learning
Shaohui Lin, Rongrong Ji, Chenqian Yan +5
cs.CVarXiv:1903.09291v12019DefaultShift: Auditing Semantic Default Shift in Accelerated Text-to-Image Models
Xuanhua Yin, Chuanzhi Xu, Shunqi Mao +2
cs.CVarXiv:2608.21784v12026Improving Computer-aided Detection using Convolutional Neural Networks and Random View Aggregation
Holger R. Roth, Le Lu, Jiamin Liu +5
cs.CVarXiv:1505.03046v22015Calibrate What You SHIP: Post-Selection Risk Control for Verifier-Guided Text-to-Image Generation
Xuanhua Yin, Shunqi Mao, Wei Guo +2
cs.CVarXiv:2608.21748v12026Geometric GAN
Jae Hyun Lim, Jong Chul Ye
stat.MLcond-mat.dis-nncs.AIarXiv:1705.02894v22017Learning to Compose Neural Networks for Question Answering
Jacob Andreas, Marcus Rohrbach, Trevor Darrell +1
cs.CLcs.CVcs.NEarXiv:1601.01705v42016Attend, Infer, Repeat: Fast Scene Understanding with Generative Models
S. M. Ali Eslami, Nicolas Heess, Theophane Weber +4
cs.CVcs.LGarXiv:1603.08575v32016f-VAEGAN-D2: A Feature Generating Framework for Any-Shot Learning
Yongqin Xian, Saurabh Sharma, Bernt Schiele +1
cs.CVarXiv:1903.10132v12019Phenaki: Variable Length Video Generation From Open Domain Textual Description
Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans +6
cs.CVcs.AIarXiv:2210.02399v12022A Multi-View Embedding Space for Modeling Internet Images, Tags, and their Semantics
Yunchao Gong, Qifa Ke, Michael Isard +1
cs.CVcs.IRcs.LGarXiv:1212.4522v22012TextCaps: a Dataset for Image Captioning with Reading Comprehension
Oleksii Sidorov, Ronghang Hu, Marcus Rohrbach +1
cs.CVcs.CLarXiv:2003.12462v22020Flowing ConvNets for Human Pose Estimation in Videos
Tomas Pfister, James Charles, Andrew Zisserman
cs.CVarXiv:1506.02897v22015Co-occurrence Feature Learning from Skeleton Data for Action Recognition and Detection with Hierarchical Aggregation
Chao Li, Qiaoyong Zhong, Di Xie +1
cs.CVarXiv:1804.06055v12018V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs
Penghao Wu, Saining Xie
cs.CVarXiv:2312.14135v22023Unite the People: Closing the Loop Between 3D and 2D Human Representations
Christoph Lassner, Javier Romero, Martin Kiefel +3
cs.CVarXiv:1701.02468v32017Convolutional neural network architecture for geometric matching
Ignacio Rocco, Relja Arandjelović, Josef Sivic
cs.CVcs.LGarXiv:1703.05593v22017Inferring and Executing Programs for Visual Reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten +4
cs.CVcs.CLcs.LGarXiv:1705.03633v12017Pretreatment DCE-MRI Resolves Response Quality Within Pathologic Endpoints in Neoadjuvant Breast Cancer
Dattatreya Kantha, Murray H. Loew
eess.IVcs.CVcs.LGarXiv:2608.22097v12026Input-Aware Dynamic Backdoor Attack
Anh Nguyen, Anh Tran
cs.CRcs.CVarXiv:2010.08138v12020End-to-end people detection in crowded scenes
Russell Stewart, Mykhaylo Andriluka
cs.CVarXiv:1506.04878v32015DIRE for Diffusion-Generated Image Detection
Zhendong Wang, Jianmin Bao, Wengang Zhou +4
cs.CVarXiv:2303.09295v12023Binary Neural Networks: A Survey
Haotong Qin, Ruihao Gong, Xianglong Liu +3
cs.NEcs.CVcs.LGarXiv:2004.03333v12020VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models
Haodong Duan, Xinyu Fang, Junming Yang +41
cs.CVarXiv:2407.11691v52024A New 2.5D Representation for Lymph Node Detection using Random Sets of Deep Convolutional Neural Network Observations
Holger R. Roth, Le Lu, Ari Seff +6
cs.CVcs.LGcs.NEarXiv:1406.2639v12014Learning a Multi-View Stereo Machine
Abhishek Kar, Christian Häne, Jitendra Malik
cs.CVarXiv:1708.05375v12017Spiking Neural Networks for Energy-Efficient Object Detection in Forward-Looking Sonar Imagery
Gwenevere Frank, Gert Cauwenberghs
cs.CVeess.SParXiv:2608.22072v12026Low-Light Image and Video Enhancement Using Deep Learning: A Survey
Chongyi Li, Chunle Guo, Linghao Han +4
cs.CVarXiv:2104.10729v32021VIG: Visual Information Gain as a Reward Signal for Multimodal Chain-of-Thought Compression
Wen Luo, Xiaohan Yi, Xiaotao Huang +1
cs.CVarXiv:2608.21883v12026Stochastic Variational Video Prediction
Mohammad Babaeizadeh, Chelsea Finn, Dumitru Erhan +2
cs.CVcs.ROarXiv:1710.11252v22017Image Captioning: Transforming Objects into Words
Simao Herdade, Armin Kappeler, Kofi Boakye +1
cs.CVcs.CLarXiv:1906.05963v22019Visual Saliency Based on Scale-Space Analysis in the Frequency Domain
Jian Li, Martin Levine, Xiangjing An +2
cs.CVarXiv:1605.01999v12016A Survey of Recent Advances in CNN-based Single Image Crowd Counting and Density Estimation
Vishwanath A. Sindagi, Vishal M. Patel
cs.CVarXiv:1707.01202v12017Learning Human-Object Interactions by Graph Parsing Neural Networks
Siyuan Qi, Wenguan Wang, Baoxiong Jia +2
cs.CVarXiv:1808.07962v12018Pointwise Convolutional Neural Networks
Binh-Son Hua, Minh-Khoi Tran, Sai-Kit Yeung
cs.CVcs.LGarXiv:1712.05245v22017