Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
2,161 to 2,220 of 18,780
Semantic-Aware Domain Generalized Segmentation
Duo Peng, Yinjie Lei, Munawar Hayat +2
cs.CVarXiv:2204.00822v12022Revealing Single Frame Bias for Video-and-Language Learning
Jie Lei, Tamara L. Berg, Mohit Bansal
cs.CVcs.AIcs.CLarXiv:2206.03428v12022Contrastive Learning for Weakly Supervised Phrase Grounding
Tanmay Gupta, Arash Vahdat, Gal Chechik +3
cs.CVcs.CLcs.LGarXiv:2006.09920v32020Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMs
Xiaofu Chen, Stella Frank, Yova Kementchedjhieva
cs.CVcs.AIarXiv:2609.09124v12026POCOVID-Net: Automatic Detection of COVID-19 From a New Lung Ultrasound Imaging Dataset (POCUS)
Jannis Born, Gabriel Brändle, Manuel Cossio +4
eess.IVcs.CVcs.LGarXiv:2004.12084v42020Single-Model and Any-Modality for Video Object Tracking
Zongwei Wu, Jilai Zheng, Xiangxuan Ren +5
cs.CVarXiv:2311.15851v32023SIDA: Social Media Image Deepfake Detection, Localization and Explanation with Large Multimodal Model
Zhenglin Huang, Jinwei Hu, Xiangtai Li +6
cs.CVcs.AIarXiv:2412.04292v32024SCOPS: Self-Supervised Co-Part Segmentation
Wei-Chih Hung, Varun Jampani, Sifei Liu +3
cs.CVarXiv:1905.01298v12019Rotation-Invariant Transformer for Point Cloud Matching
Hao Yu, Zheng Qin, Ji Hou +4
cs.CVarXiv:2303.08231v32023Drivable 3D Gaussian Avatars
Wojciech Zielonka, Timur Bagautdinov, Shunsuke Saito +3
cs.CVarXiv:2311.08581v22023Parallel Multi Channel Convolution using General Matrix Multiplication
Aravind Vasudevan, Andrew Anderson, David Gregg
cs.CVcs.PFarXiv:1704.04428v22017Image Synthesis with Adversarial Networks: a Comprehensive Survey and Case Studies
Pourya Shamsolmoali, Masoumeh Zareapoor, Eric Granger +4
cs.CVeess.IVarXiv:2012.13736v12020Deep Learning-based Face Super-Resolution: A Survey
Junjun Jiang, Chenyang Wang, Xianming Liu +1
cs.CVarXiv:2101.03749v22021Haze Visibility Enhancement: A Survey and Quantitative Benchmarking
Yu Li, Shaodi You, Michael S. Brown +1
cs.CVarXiv:1607.06235v12016Bayesian Image Quality Transfer with CNNs: Exploring Uncertainty in dMRI Super-Resolution
Ryutaro Tanno, Daniel E. Worrall, Aurobrata Ghosh +4
cs.CVarXiv:1705.00664v22017Driving Style Analysis Using Primitive Driving Patterns With Bayesian Nonparametric Approaches
Wenshuo Wang, Junqiang Xi, Ding Zhao
cs.CVarXiv:1708.08986v12017Fast convolutional neural networks on FPGAs with hls4ml
Thea Aarrestad, Vladimir Loncar, Nicolò Ghielmetti +17
cs.LGcs.CVhep-exarXiv:2101.05108v22021Comparing the Performance of L*A*B* and HSV Color Spaces with Respect to Color Image Segmentation
Dibya Jyoti Bora, Anil Kumar Gupta, Fayaz Ahmad Khan
cs.CVarXiv:1506.01472v12015Foley Music: Learning to Generate Music from Videos
Chuang Gan, Deng Huang, Peihao Chen +2
cs.CVcs.LGcs.SDarXiv:2007.10984v12020HVPR: Hybrid Voxel-Point Representation for Single-stage 3D Object Detection
Jongyoun Noh, Sanghoon Lee, Bumsub Ham
cs.CVarXiv:2104.00902v12021Hi-FLoop: Hierarchical State-Feedback Loops for Multi-Timescale World Modeling
Rx Fan, Zhan H
cs.CVcs.AIarXiv:2609.08796v12026UniPose: Unified Human Pose Estimation in Single Images and Videos
Bruno Artacho, Andreas Savakis
cs.CVarXiv:2001.08095v12020Stereo obstacle detection for unmanned surface vehicles by IMU-assisted semantic segmentation
Borja Bovcon, Rok Mandeljc, Janez Perš +1
cs.ROcs.CVarXiv:1802.07956v12018A Bottom-up Approach for Pancreas Segmentation using Cascaded Superpixels and (Deep) Image Patch Labeling
Amal Farag, Le Lu, Holger R. Roth +3
cs.CVarXiv:1505.06236v22015Unsupervised Person Re-identification by Deep Asymmetric Metric Embedding
Hong-Xing Yu, Ancong Wu, Wei-Shi Zheng
cs.CVarXiv:1901.10177v12019Deconfounded Image Captioning: A Causal Retrospect
Xu Yang, Hanwang Zhang, Jianfei Cai
cs.CVarXiv:2003.03923v22020JSENet: Joint Semantic Segmentation and Edge Detection Network for 3D Point Clouds
Zeyu Hu, Mingmin Zhen, Xuyang Bai +2
cs.CVarXiv:2007.06888v12020Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics
Ruibo Ming, Lei Sun, Deheng Zhang +8
cs.CVcs.AIarXiv:2609.08755v12026What's Cookin'? Interpreting Cooking Videos using Text, Speech and Vision
Jonathan Malmaud, Jonathan Huang, Vivek Rathod +3
cs.CLcs.CVcs.IRarXiv:1503.01558v32015Deep Graph-Convolutional Image Denoising
Diego Valsesia, Giulia Fracastoro, Enrico Magli
eess.IVcs.CVcs.LGarXiv:1907.08448v12019Low-rank Kernel Learning for Graph-based Clustering
Zhao Kang, Liangjian Wen, Wenyu Chen +1
cs.LGcs.CVstat.MLarXiv:1903.05962v12019CausalChapter: Improving Long-Video Chaptering with Interventional Dependency Modeling
Xinran Duan, Guozhang Li, Yaoyao Zhong +3
cs.CVcs.AIarXiv:2609.08686v12026UcoSLAM: Simultaneous Localization and Mapping by Fusion of KeyPoints and Squared Planar Markers
Rafael Munoz-Salinas, Rafael Medina-Carnicer
cs.CVarXiv:1902.03729v12019Clustering with Multi-Layer Graphs: A Spectral Perspective
Xiaowen Dong, Pascal Frossard, Pierre Vandergheynst +1
cs.LGcs.CVcs.SIarXiv:1106.2233v12011Segment Any Point Cloud Sequences by Distilling Vision Foundation Models
Youquan Liu, Lingdong Kong, Jun Cen +5
cs.CVcs.LGcs.ROarXiv:2306.09347v22023DeepSOCIAL: Social Distancing Monitoring and Infection Risk Assessment in COVID-19 Pandemic
Mahdi Rezaei, Mohsen Azarmi
cs.CVcs.LGeess.IVarXiv:2008.11672v32020COVID-19 Chest CT Image Segmentation -- A Deep Convolutional Neural Network Solution
Qingsen Yan, Bo Wang, Dong Gong +7
eess.IVcs.CVcs.LGarXiv:2004.10987v22020AxonDeepSeg: automatic axon and myelin segmentation from microscopy data using convolutional neural networks
Aldo Zaimi, Maxime Wabartha, Victor Herman +3
cs.CVarXiv:1711.01004v22017Deep Generative Adversarial Networks for Compressed Sensing Automates MRI
Morteza Mardani, Enhao Gong, Joseph Y. Cheng +8
cs.CVcs.LGstat.MLarXiv:1706.00051v12017Interactive Visual Grounding of Referring Expressions for Human-Robot Interaction
Mohit Shridhar, David Hsu
cs.ROcs.CLcs.CVarXiv:1806.03831v12018Context-Aware Mixup for Domain Adaptive Semantic Segmentation
Qianyu Zhou, Zhengyang Feng, Qiqi Gu +5
cs.CVarXiv:2108.03557v320216-DoF Pose Estimation of Household Objects for Robotic Manipulation: An Accessible Dataset and Benchmark
Stephen Tyree, Jonathan Tremblay, Thang To +4
cs.ROcs.CVarXiv:2203.05701v22022TriCCOT: Tri-part Convolutional Conformal Transformer for Onboard Space Object Detection
Adrien Dorise, Marjorie Bellizzi, Julia Cohen +1
cs.CVcs.AIarXiv:2609.08659v12026Infinite Feature Selection: A Graph-based Feature Filtering Approach
Giorgio Roffo, Simone Melzi, Umberto Castellani +2
cs.CVcs.LGstat.MLarXiv:2006.08184v12020Active Fire Detection in Landsat-8 Imagery: a Large-Scale Dataset and a Deep-Learning Study
Gabriel Henrique de Almeida Pereira, André Minoro Fusioka, Bogdan Tomoyuki Nassu +1
cs.CVcs.LGarXiv:2101.03409v22021A Unified Deep Neural Network for Speaker and Language Recognition
Fred Richardson, Douglas Reynolds, Najim Dehak
cs.CLcs.CVcs.LGarXiv:1504.00923v12015Fully Convolutional Network for Automatic Road Extraction from Satellite Imagery
Alexander V. Buslaev, Selim S. Seferbekov, Vladimir I. Iglovikov +1
cs.CVarXiv:1806.05182v22018From Where to How: Continuous 4D Interaction Forecasting from Egocentric Video
Qiaohui Chu, Haoyu Zhang, Meng Liu +3
cs.CVcs.AIarXiv:2609.08636v12026Spatially-Adaptive Image Restoration using Distortion-Guided Networks
Kuldeep Purohit, Maitreya Suin, A. N. Rajagopalan +1
cs.CVcs.LGarXiv:2108.08617v12021GiraffeDet: A Heavy-Neck Paradigm for Object Detection
Yiqi Jiang, Zhiyu Tan, Junyan Wang +3
cs.CVarXiv:2202.04256v22022Improving the Efficiency and Robustness of Deepfakes Detection through Precise Geometric Features
Zekun Sun, Yujie Han, Zeyu Hua +2
cs.CVarXiv:2104.04480v12021Fast Sparse ConvNets
Erich Elsen, Marat Dukhan, Trevor Gale +1
cs.CVarXiv:1911.09723v12019Weakly Supervised Contrastive Learning
Mingkai Zheng, Fei Wang, Shan You +4
cs.CVarXiv:2110.04770v12021Mixed Neural Voxels for Fast Multi-view Video Synthesis
Feng Wang, Sinan Tan, Xinghang Li +3
cs.CVarXiv:2212.00190v22022Label-driven weakly-supervised learning for multimodal deformable image registration
Yipeng Hu, Marc Modat, Eli Gibson +7
cs.CVcs.LGarXiv:1711.01666v22017Multi-level Attention network using text, audio and video for Depression Prediction
Anupama Ray, Siddharth Kumar, Rutvik Reddy +2
cs.CVeess.ASarXiv:1909.01417v12019A Multi-Scale CNN and Curriculum Learning Strategy for Mammogram Classification
William Lotter, Greg Sorensen, David Cox
cs.CVarXiv:1707.06978v12017Learning monocular depth estimation with unsupervised trinocular assumptions
Matteo Poggi, Fabio Tosi, Stefano Mattoccia
cs.CVarXiv:1808.01606v12018Camera Lens Super-Resolution
Chang Chen, Zhiwei Xiong, Xinmei Tian +2
cs.CVarXiv:1904.03378v12019MiniSeg: An Extremely Minimum Network for Efficient COVID-19 Segmentation
Yu Qiu, Yun Liu, Shijie Li +1
cs.CVarXiv:2004.09750v32020