Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

2,221 to 2,280 of 18,817

  1. SCOPS: Self-Supervised Co-Part Segmentation

    Wei-Chih Hung, Varun Jampani, Sifei Liu +3

    cs.CVarXiv:1905.01298v12019
  2. Rotation-Invariant Transformer for Point Cloud Matching

    Hao Yu, Zheng Qin, Ji Hou +4

    cs.CVarXiv:2303.08231v32023
  3. Drivable 3D Gaussian Avatars

    Wojciech Zielonka, Timur Bagautdinov, Shunsuke Saito +3

    cs.CVarXiv:2311.08581v22023
  4. Parallel Multi Channel Convolution using General Matrix Multiplication

    Aravind Vasudevan, Andrew Anderson, David Gregg

    cs.CVcs.PFarXiv:1704.04428v22017
  5. Image Synthesis with Adversarial Networks: a Comprehensive Survey and Case Studies

    Pourya Shamsolmoali, Masoumeh Zareapoor, Eric Granger +4

    cs.CVeess.IVarXiv:2012.13736v12020
  6. Deep Learning-based Face Super-Resolution: A Survey

    Junjun Jiang, Chenyang Wang, Xianming Liu +1

    cs.CVarXiv:2101.03749v22021
  7. Haze Visibility Enhancement: A Survey and Quantitative Benchmarking

    Yu Li, Shaodi You, Michael S. Brown +1

    cs.CVarXiv:1607.06235v12016
  8. Bayesian Image Quality Transfer with CNNs: Exploring Uncertainty in dMRI Super-Resolution

    Ryutaro Tanno, Daniel E. Worrall, Aurobrata Ghosh +4

    cs.CVarXiv:1705.00664v22017
  9. Driving Style Analysis Using Primitive Driving Patterns With Bayesian Nonparametric Approaches

    Wenshuo Wang, Junqiang Xi, Ding Zhao

    cs.CVarXiv:1708.08986v12017
  10. Fast convolutional neural networks on FPGAs with hls4ml

    Thea Aarrestad, Vladimir Loncar, Nicolò Ghielmetti +17

    cs.LGcs.CVhep-exarXiv:2101.05108v22021
  11. Comparing the Performance of L*A*B* and HSV Color Spaces with Respect to Color Image Segmentation

    Dibya Jyoti Bora, Anil Kumar Gupta, Fayaz Ahmad Khan

    cs.CVarXiv:1506.01472v12015
  12. Foley Music: Learning to Generate Music from Videos

    Chuang Gan, Deng Huang, Peihao Chen +2

    cs.CVcs.LGcs.SDarXiv:2007.10984v12020
  13. HVPR: Hybrid Voxel-Point Representation for Single-stage 3D Object Detection

    Jongyoun Noh, Sanghoon Lee, Bumsub Ham

    cs.CVarXiv:2104.00902v12021
  14. Hi-FLoop: Hierarchical State-Feedback Loops for Multi-Timescale World Modeling

    Rx Fan, Zhan H

    cs.CVcs.AIarXiv:2609.08796v12026
  15. UniPose: Unified Human Pose Estimation in Single Images and Videos

    Bruno Artacho, Andreas Savakis

    cs.CVarXiv:2001.08095v12020
  16. Stereo obstacle detection for unmanned surface vehicles by IMU-assisted semantic segmentation

    Borja Bovcon, Rok Mandeljc, Janez Perš +1

    cs.ROcs.CVarXiv:1802.07956v12018
  17. A Bottom-up Approach for Pancreas Segmentation using Cascaded Superpixels and (Deep) Image Patch Labeling

    Amal Farag, Le Lu, Holger R. Roth +3

    cs.CVarXiv:1505.06236v22015
  18. Unsupervised Person Re-identification by Deep Asymmetric Metric Embedding

    Hong-Xing Yu, Ancong Wu, Wei-Shi Zheng

    cs.CVarXiv:1901.10177v12019
  19. Deconfounded Image Captioning: A Causal Retrospect

    Xu Yang, Hanwang Zhang, Jianfei Cai

    cs.CVarXiv:2003.03923v22020
  20. JSENet: Joint Semantic Segmentation and Edge Detection Network for 3D Point Clouds

    Zeyu Hu, Mingmin Zhen, Xuyang Bai +2

    cs.CVarXiv:2007.06888v12020
  21. Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics

    Ruibo Ming, Lei Sun, Deheng Zhang +8

    cs.CVcs.AIarXiv:2609.08755v12026
  22. What's Cookin'? Interpreting Cooking Videos using Text, Speech and Vision

    Jonathan Malmaud, Jonathan Huang, Vivek Rathod +3

    cs.CLcs.CVcs.IRarXiv:1503.01558v32015
  23. Deep Graph-Convolutional Image Denoising

    Diego Valsesia, Giulia Fracastoro, Enrico Magli

    eess.IVcs.CVcs.LGarXiv:1907.08448v12019
  24. Low-rank Kernel Learning for Graph-based Clustering

    Zhao Kang, Liangjian Wen, Wenyu Chen +1

    cs.LGcs.CVstat.MLarXiv:1903.05962v12019
  25. CausalChapter: Improving Long-Video Chaptering with Interventional Dependency Modeling

    Xinran Duan, Guozhang Li, Yaoyao Zhong +3

    cs.CVcs.AIarXiv:2609.08686v12026
  26. UcoSLAM: Simultaneous Localization and Mapping by Fusion of KeyPoints and Squared Planar Markers

    Rafael Munoz-Salinas, Rafael Medina-Carnicer

    cs.CVarXiv:1902.03729v12019
  27. Clustering with Multi-Layer Graphs: A Spectral Perspective

    Xiaowen Dong, Pascal Frossard, Pierre Vandergheynst +1

    cs.LGcs.CVcs.SIarXiv:1106.2233v12011
  28. Segment Any Point Cloud Sequences by Distilling Vision Foundation Models

    Youquan Liu, Lingdong Kong, Jun Cen +5

    cs.CVcs.LGcs.ROarXiv:2306.09347v22023
  29. DeepSOCIAL: Social Distancing Monitoring and Infection Risk Assessment in COVID-19 Pandemic

    Mahdi Rezaei, Mohsen Azarmi

    cs.CVcs.LGeess.IVarXiv:2008.11672v32020
  30. COVID-19 Chest CT Image Segmentation -- A Deep Convolutional Neural Network Solution

    Qingsen Yan, Bo Wang, Dong Gong +7

    eess.IVcs.CVcs.LGarXiv:2004.10987v22020
  31. AxonDeepSeg: automatic axon and myelin segmentation from microscopy data using convolutional neural networks

    Aldo Zaimi, Maxime Wabartha, Victor Herman +3

    cs.CVarXiv:1711.01004v22017
  32. Deep Generative Adversarial Networks for Compressed Sensing Automates MRI

    Morteza Mardani, Enhao Gong, Joseph Y. Cheng +8

    cs.CVcs.LGstat.MLarXiv:1706.00051v12017
  33. Interactive Visual Grounding of Referring Expressions for Human-Robot Interaction

    Mohit Shridhar, David Hsu

    cs.ROcs.CLcs.CVarXiv:1806.03831v12018
  34. Context-Aware Mixup for Domain Adaptive Semantic Segmentation

    Qianyu Zhou, Zhengyang Feng, Qiqi Gu +5

    cs.CVarXiv:2108.03557v32021
  35. 6-DoF Pose Estimation of Household Objects for Robotic Manipulation: An Accessible Dataset and Benchmark

    Stephen Tyree, Jonathan Tremblay, Thang To +4

    cs.ROcs.CVarXiv:2203.05701v22022
  36. TriCCOT: Tri-part Convolutional Conformal Transformer for Onboard Space Object Detection

    Adrien Dorise, Marjorie Bellizzi, Julia Cohen +1

    cs.CVcs.AIarXiv:2609.08659v12026
  37. Infinite Feature Selection: A Graph-based Feature Filtering Approach

    Giorgio Roffo, Simone Melzi, Umberto Castellani +2

    cs.CVcs.LGstat.MLarXiv:2006.08184v12020
  38. Active Fire Detection in Landsat-8 Imagery: a Large-Scale Dataset and a Deep-Learning Study

    Gabriel Henrique de Almeida Pereira, André Minoro Fusioka, Bogdan Tomoyuki Nassu +1

    cs.CVcs.LGarXiv:2101.03409v22021
  39. A Unified Deep Neural Network for Speaker and Language Recognition

    Fred Richardson, Douglas Reynolds, Najim Dehak

    cs.CLcs.CVcs.LGarXiv:1504.00923v12015
  40. Fully Convolutional Network for Automatic Road Extraction from Satellite Imagery

    Alexander V. Buslaev, Selim S. Seferbekov, Vladimir I. Iglovikov +1

    cs.CVarXiv:1806.05182v22018
  41. From Where to How: Continuous 4D Interaction Forecasting from Egocentric Video

    Qiaohui Chu, Haoyu Zhang, Meng Liu +3

    cs.CVcs.AIarXiv:2609.08636v12026
  42. Spatially-Adaptive Image Restoration using Distortion-Guided Networks

    Kuldeep Purohit, Maitreya Suin, A. N. Rajagopalan +1

    cs.CVcs.LGarXiv:2108.08617v12021
  43. GiraffeDet: A Heavy-Neck Paradigm for Object Detection

    Yiqi Jiang, Zhiyu Tan, Junyan Wang +3

    cs.CVarXiv:2202.04256v22022
  44. Improving the Efficiency and Robustness of Deepfakes Detection through Precise Geometric Features

    Zekun Sun, Yujie Han, Zeyu Hua +2

    cs.CVarXiv:2104.04480v12021
  45. Fast Sparse ConvNets

    Erich Elsen, Marat Dukhan, Trevor Gale +1

    cs.CVarXiv:1911.09723v12019
  46. Weakly Supervised Contrastive Learning

    Mingkai Zheng, Fei Wang, Shan You +4

    cs.CVarXiv:2110.04770v12021
  47. Mixed Neural Voxels for Fast Multi-view Video Synthesis

    Feng Wang, Sinan Tan, Xinghang Li +3

    cs.CVarXiv:2212.00190v22022
  48. Label-driven weakly-supervised learning for multimodal deformable image registration

    Yipeng Hu, Marc Modat, Eli Gibson +7

    cs.CVcs.LGarXiv:1711.01666v22017
  49. Multi-level Attention network using text, audio and video for Depression Prediction

    Anupama Ray, Siddharth Kumar, Rutvik Reddy +2

    cs.CVeess.ASarXiv:1909.01417v12019
  50. A Multi-Scale CNN and Curriculum Learning Strategy for Mammogram Classification

    William Lotter, Greg Sorensen, David Cox

    cs.CVarXiv:1707.06978v12017
  51. Learning monocular depth estimation with unsupervised trinocular assumptions

    Matteo Poggi, Fabio Tosi, Stefano Mattoccia

    cs.CVarXiv:1808.01606v12018
  52. Camera Lens Super-Resolution

    Chang Chen, Zhiwei Xiong, Xinmei Tian +2

    cs.CVarXiv:1904.03378v12019
  53. MiniSeg: An Extremely Minimum Network for Efficient COVID-19 Segmentation

    Yu Qiu, Yun Liu, Shijie Li +1

    cs.CVarXiv:2004.09750v32020
  54. Two-stage framework for optic disc localization and glaucoma classification in retinal fundus images using deep learning

    Muhammad Naseer Bajwa, Muhammad Imran Malik, Shoaib Ahmed Siddiqui +4

    cs.CVcs.LGeess.IVarXiv:2005.14284v12020
  55. The Visual Insensitivity Gap: Diagnosing When Vision-Language Models Fail to Use Visual Evidence

    Genpei Zhang

    cs.CVcs.CLcs.LGarXiv:2609.00868v12026
  56. Seeing Around Street Corners: Non-Line-of-Sight Detection and Tracking In-the-Wild Using Doppler Radar

    Nicolas Scheiner, Florian Kraus, Fangyin Wei +8

    cs.CVeess.IVarXiv:1912.06613v22019
  57. MegaPortraits: One-shot Megapixel Neural Head Avatars

    Nikita Drobyshev, Jenya Chelishev, Taras Khakhulin +3

    cs.CVarXiv:2207.07621v22022
  58. SynthRCT: Scalable Conditional Deformation Synthesis for Synthetic Repeat CT Generation

    Tomas Guija-Valiente, Blanca Rodriguez-Gonzalez, Norberto Malpica

    cs.CVcs.AIcs.LGarXiv:2609.08627v12026
  59. POSEidon: Face-from-Depth for Driver Pose Estimation

    Guido Borghi, Marco Venturelli, Roberto Vezzani +1

    cs.CVarXiv:1611.10195v32016
  60. Audio-Visual Speech Recognition With A Hybrid CTC/Attention Architecture

    Stavros Petridis, Themos Stafylakis, Pingchuan Ma +2

    cs.CVarXiv:1810.00108v12018