Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

2,161 to 2,220 of 18,780

  1. Semantic-Aware Domain Generalized Segmentation

    Duo Peng, Yinjie Lei, Munawar Hayat +2

    cs.CVarXiv:2204.00822v12022
  2. Revealing Single Frame Bias for Video-and-Language Learning

    Jie Lei, Tamara L. Berg, Mohit Bansal

    cs.CVcs.AIcs.CLarXiv:2206.03428v12022
  3. Contrastive Learning for Weakly Supervised Phrase Grounding

    Tanmay Gupta, Arash Vahdat, Gal Chechik +3

    cs.CVcs.CLcs.LGarXiv:2006.09920v32020
  4. Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMs

    Xiaofu Chen, Stella Frank, Yova Kementchedjhieva

    cs.CVcs.AIarXiv:2609.09124v12026
  5. POCOVID-Net: Automatic Detection of COVID-19 From a New Lung Ultrasound Imaging Dataset (POCUS)

    Jannis Born, Gabriel Brändle, Manuel Cossio +4

    eess.IVcs.CVcs.LGarXiv:2004.12084v42020
  6. Single-Model and Any-Modality for Video Object Tracking

    Zongwei Wu, Jilai Zheng, Xiangxuan Ren +5

    cs.CVarXiv:2311.15851v32023
  7. SIDA: Social Media Image Deepfake Detection, Localization and Explanation with Large Multimodal Model

    Zhenglin Huang, Jinwei Hu, Xiangtai Li +6

    cs.CVcs.AIarXiv:2412.04292v32024
  8. SCOPS: Self-Supervised Co-Part Segmentation

    Wei-Chih Hung, Varun Jampani, Sifei Liu +3

    cs.CVarXiv:1905.01298v12019
  9. Rotation-Invariant Transformer for Point Cloud Matching

    Hao Yu, Zheng Qin, Ji Hou +4

    cs.CVarXiv:2303.08231v32023
  10. Drivable 3D Gaussian Avatars

    Wojciech Zielonka, Timur Bagautdinov, Shunsuke Saito +3

    cs.CVarXiv:2311.08581v22023
  11. Parallel Multi Channel Convolution using General Matrix Multiplication

    Aravind Vasudevan, Andrew Anderson, David Gregg

    cs.CVcs.PFarXiv:1704.04428v22017
  12. Image Synthesis with Adversarial Networks: a Comprehensive Survey and Case Studies

    Pourya Shamsolmoali, Masoumeh Zareapoor, Eric Granger +4

    cs.CVeess.IVarXiv:2012.13736v12020
  13. Deep Learning-based Face Super-Resolution: A Survey

    Junjun Jiang, Chenyang Wang, Xianming Liu +1

    cs.CVarXiv:2101.03749v22021
  14. Haze Visibility Enhancement: A Survey and Quantitative Benchmarking

    Yu Li, Shaodi You, Michael S. Brown +1

    cs.CVarXiv:1607.06235v12016
  15. Bayesian Image Quality Transfer with CNNs: Exploring Uncertainty in dMRI Super-Resolution

    Ryutaro Tanno, Daniel E. Worrall, Aurobrata Ghosh +4

    cs.CVarXiv:1705.00664v22017
  16. Driving Style Analysis Using Primitive Driving Patterns With Bayesian Nonparametric Approaches

    Wenshuo Wang, Junqiang Xi, Ding Zhao

    cs.CVarXiv:1708.08986v12017
  17. Fast convolutional neural networks on FPGAs with hls4ml

    Thea Aarrestad, Vladimir Loncar, Nicolò Ghielmetti +17

    cs.LGcs.CVhep-exarXiv:2101.05108v22021
  18. Comparing the Performance of L*A*B* and HSV Color Spaces with Respect to Color Image Segmentation

    Dibya Jyoti Bora, Anil Kumar Gupta, Fayaz Ahmad Khan

    cs.CVarXiv:1506.01472v12015
  19. Foley Music: Learning to Generate Music from Videos

    Chuang Gan, Deng Huang, Peihao Chen +2

    cs.CVcs.LGcs.SDarXiv:2007.10984v12020
  20. HVPR: Hybrid Voxel-Point Representation for Single-stage 3D Object Detection

    Jongyoun Noh, Sanghoon Lee, Bumsub Ham

    cs.CVarXiv:2104.00902v12021
  21. Hi-FLoop: Hierarchical State-Feedback Loops for Multi-Timescale World Modeling

    Rx Fan, Zhan H

    cs.CVcs.AIarXiv:2609.08796v12026
  22. UniPose: Unified Human Pose Estimation in Single Images and Videos

    Bruno Artacho, Andreas Savakis

    cs.CVarXiv:2001.08095v12020
  23. Stereo obstacle detection for unmanned surface vehicles by IMU-assisted semantic segmentation

    Borja Bovcon, Rok Mandeljc, Janez Perš +1

    cs.ROcs.CVarXiv:1802.07956v12018
  24. A Bottom-up Approach for Pancreas Segmentation using Cascaded Superpixels and (Deep) Image Patch Labeling

    Amal Farag, Le Lu, Holger R. Roth +3

    cs.CVarXiv:1505.06236v22015
  25. Unsupervised Person Re-identification by Deep Asymmetric Metric Embedding

    Hong-Xing Yu, Ancong Wu, Wei-Shi Zheng

    cs.CVarXiv:1901.10177v12019
  26. Deconfounded Image Captioning: A Causal Retrospect

    Xu Yang, Hanwang Zhang, Jianfei Cai

    cs.CVarXiv:2003.03923v22020
  27. JSENet: Joint Semantic Segmentation and Edge Detection Network for 3D Point Clouds

    Zeyu Hu, Mingmin Zhen, Xuyang Bai +2

    cs.CVarXiv:2007.06888v12020
  28. Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics

    Ruibo Ming, Lei Sun, Deheng Zhang +8

    cs.CVcs.AIarXiv:2609.08755v12026
  29. What's Cookin'? Interpreting Cooking Videos using Text, Speech and Vision

    Jonathan Malmaud, Jonathan Huang, Vivek Rathod +3

    cs.CLcs.CVcs.IRarXiv:1503.01558v32015
  30. Deep Graph-Convolutional Image Denoising

    Diego Valsesia, Giulia Fracastoro, Enrico Magli

    eess.IVcs.CVcs.LGarXiv:1907.08448v12019
  31. Low-rank Kernel Learning for Graph-based Clustering

    Zhao Kang, Liangjian Wen, Wenyu Chen +1

    cs.LGcs.CVstat.MLarXiv:1903.05962v12019
  32. CausalChapter: Improving Long-Video Chaptering with Interventional Dependency Modeling

    Xinran Duan, Guozhang Li, Yaoyao Zhong +3

    cs.CVcs.AIarXiv:2609.08686v12026
  33. UcoSLAM: Simultaneous Localization and Mapping by Fusion of KeyPoints and Squared Planar Markers

    Rafael Munoz-Salinas, Rafael Medina-Carnicer

    cs.CVarXiv:1902.03729v12019
  34. Clustering with Multi-Layer Graphs: A Spectral Perspective

    Xiaowen Dong, Pascal Frossard, Pierre Vandergheynst +1

    cs.LGcs.CVcs.SIarXiv:1106.2233v12011
  35. Segment Any Point Cloud Sequences by Distilling Vision Foundation Models

    Youquan Liu, Lingdong Kong, Jun Cen +5

    cs.CVcs.LGcs.ROarXiv:2306.09347v22023
  36. DeepSOCIAL: Social Distancing Monitoring and Infection Risk Assessment in COVID-19 Pandemic

    Mahdi Rezaei, Mohsen Azarmi

    cs.CVcs.LGeess.IVarXiv:2008.11672v32020
  37. COVID-19 Chest CT Image Segmentation -- A Deep Convolutional Neural Network Solution

    Qingsen Yan, Bo Wang, Dong Gong +7

    eess.IVcs.CVcs.LGarXiv:2004.10987v22020
  38. AxonDeepSeg: automatic axon and myelin segmentation from microscopy data using convolutional neural networks

    Aldo Zaimi, Maxime Wabartha, Victor Herman +3

    cs.CVarXiv:1711.01004v22017
  39. Deep Generative Adversarial Networks for Compressed Sensing Automates MRI

    Morteza Mardani, Enhao Gong, Joseph Y. Cheng +8

    cs.CVcs.LGstat.MLarXiv:1706.00051v12017
  40. Interactive Visual Grounding of Referring Expressions for Human-Robot Interaction

    Mohit Shridhar, David Hsu

    cs.ROcs.CLcs.CVarXiv:1806.03831v12018
  41. Context-Aware Mixup for Domain Adaptive Semantic Segmentation

    Qianyu Zhou, Zhengyang Feng, Qiqi Gu +5

    cs.CVarXiv:2108.03557v32021
  42. 6-DoF Pose Estimation of Household Objects for Robotic Manipulation: An Accessible Dataset and Benchmark

    Stephen Tyree, Jonathan Tremblay, Thang To +4

    cs.ROcs.CVarXiv:2203.05701v22022
  43. TriCCOT: Tri-part Convolutional Conformal Transformer for Onboard Space Object Detection

    Adrien Dorise, Marjorie Bellizzi, Julia Cohen +1

    cs.CVcs.AIarXiv:2609.08659v12026
  44. Infinite Feature Selection: A Graph-based Feature Filtering Approach

    Giorgio Roffo, Simone Melzi, Umberto Castellani +2

    cs.CVcs.LGstat.MLarXiv:2006.08184v12020
  45. Active Fire Detection in Landsat-8 Imagery: a Large-Scale Dataset and a Deep-Learning Study

    Gabriel Henrique de Almeida Pereira, André Minoro Fusioka, Bogdan Tomoyuki Nassu +1

    cs.CVcs.LGarXiv:2101.03409v22021
  46. A Unified Deep Neural Network for Speaker and Language Recognition

    Fred Richardson, Douglas Reynolds, Najim Dehak

    cs.CLcs.CVcs.LGarXiv:1504.00923v12015
  47. Fully Convolutional Network for Automatic Road Extraction from Satellite Imagery

    Alexander V. Buslaev, Selim S. Seferbekov, Vladimir I. Iglovikov +1

    cs.CVarXiv:1806.05182v22018
  48. From Where to How: Continuous 4D Interaction Forecasting from Egocentric Video

    Qiaohui Chu, Haoyu Zhang, Meng Liu +3

    cs.CVcs.AIarXiv:2609.08636v12026
  49. Spatially-Adaptive Image Restoration using Distortion-Guided Networks

    Kuldeep Purohit, Maitreya Suin, A. N. Rajagopalan +1

    cs.CVcs.LGarXiv:2108.08617v12021
  50. GiraffeDet: A Heavy-Neck Paradigm for Object Detection

    Yiqi Jiang, Zhiyu Tan, Junyan Wang +3

    cs.CVarXiv:2202.04256v22022
  51. Improving the Efficiency and Robustness of Deepfakes Detection through Precise Geometric Features

    Zekun Sun, Yujie Han, Zeyu Hua +2

    cs.CVarXiv:2104.04480v12021
  52. Fast Sparse ConvNets

    Erich Elsen, Marat Dukhan, Trevor Gale +1

    cs.CVarXiv:1911.09723v12019
  53. Weakly Supervised Contrastive Learning

    Mingkai Zheng, Fei Wang, Shan You +4

    cs.CVarXiv:2110.04770v12021
  54. Mixed Neural Voxels for Fast Multi-view Video Synthesis

    Feng Wang, Sinan Tan, Xinghang Li +3

    cs.CVarXiv:2212.00190v22022
  55. Label-driven weakly-supervised learning for multimodal deformable image registration

    Yipeng Hu, Marc Modat, Eli Gibson +7

    cs.CVcs.LGarXiv:1711.01666v22017
  56. Multi-level Attention network using text, audio and video for Depression Prediction

    Anupama Ray, Siddharth Kumar, Rutvik Reddy +2

    cs.CVeess.ASarXiv:1909.01417v12019
  57. A Multi-Scale CNN and Curriculum Learning Strategy for Mammogram Classification

    William Lotter, Greg Sorensen, David Cox

    cs.CVarXiv:1707.06978v12017
  58. Learning monocular depth estimation with unsupervised trinocular assumptions

    Matteo Poggi, Fabio Tosi, Stefano Mattoccia

    cs.CVarXiv:1808.01606v12018
  59. Camera Lens Super-Resolution

    Chang Chen, Zhiwei Xiong, Xinmei Tian +2

    cs.CVarXiv:1904.03378v12019
  60. MiniSeg: An Extremely Minimum Network for Efficient COVID-19 Segmentation

    Yu Qiu, Yun Liu, Shijie Li +1

    cs.CVarXiv:2004.09750v32020