Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

9,001 to 9,060 of 18,866

  1. SEWA DB: A Rich Database for Audio-Visual Emotion and Sentiment Research in the Wild

    Jean Kossaifi, Robert Walecki, Yannis Panagakis +10

    cs.HCcs.AIcs.CVarXiv:1901.02839v22019
  2. Benchmarking Self-Supervised Learning on Diverse Pathology Datasets

    Mingu Kang, Heon Song, Seonwook Park +2

    cs.CVcs.LGarXiv:2212.04690v22022
  3. Multimodal Trajectory Prediction Conditioned on Lane-Graph Traversals

    Nachiket Deo, Eric M. Wolff, Oscar Beijbom

    cs.CVcs.ROarXiv:2106.15004v22021
  4. Night-to-Day Image Translation for Retrieval-based Localization

    Asha Anoosheh, Torsten Sattler, Radu Timofte +2

    cs.CVarXiv:1809.09767v22018
  5. Modular Interactive Video Object Segmentation: Interaction-to-Mask, Propagation and Difference-Aware Fusion

    Ho Kei Cheng, Yu-Wing Tai, Chi-Keung Tang

    cs.CVarXiv:2103.07941v32021
  6. RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot

    Hao-Shu Fang, Hongjie Fang, Zhenyu Tang +5

    cs.ROcs.AIcs.CVarXiv:2307.00595v22023
  7. A Large Dataset of Object Scans

    Sungjoon Choi, Qian-Yi Zhou, Stephen Miller +1

    cs.CVcs.GRarXiv:1602.02481v32016
  8. Instant-Teaching: An End-to-End Semi-Supervised Object Detection Framework

    Qiang Zhou, Chaohui Yu, Zhibin Wang +2

    cs.CVcs.AIarXiv:2103.11402v12021
  9. Material Based Object Tracking in Hyperspectral Videos: Benchmark and Algorithms

    Fengchao Xiong, Jun Zhou, Yuntao Qian

    cs.CVcs.AIarXiv:1812.04179v52018
  10. DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language Models

    Ge Zheng, Bin Yang, Jiajin Tang +2

    cs.CVcs.CLarXiv:2310.16436v22023
  11. Beyond Fluency: A Rubric-Based Benchmark for Evaluating Saudi Dialect and Cultural Competence in Large Language Models

    Ghassan Al-Sumaidaee, Sajjad Abdoli, Ahmed Rashad +1

    cs.CLcs.AIcs.CVarXiv:2608.29990v12026
  12. xMUDA: Cross-Modal Unsupervised Domain Adaptation for 3D Semantic Segmentation

    Maximilian Jaritz, Tuan-Hung Vu, Raoul de Charette +2

    cs.CVarXiv:1911.12676v22019
  13. Real-time Hand Gesture Detection and Classification Using Convolutional Neural Networks

    Okan Köpüklü, Ahmet Gunduz, Neslihan Kose +1

    cs.CVcs.AIarXiv:1901.10323v32019
  14. ViTAA: Visual-Textual Attributes Alignment in Person Search by Natural Language

    Zhe Wang, Zhiyuan Fang, Jun Wang +1

    cs.CVarXiv:2005.07327v22020
  15. LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control

    Jianzhu Guo, Dingyun Zhang, Xiaoqiang Liu +4

    cs.CVarXiv:2407.03168v22024
  16. Instance-aware Image Colorization

    Jheng-Wei Su, Hung-Kuo Chu, Jia-Bin Huang

    cs.CVarXiv:2005.10825v12020
  17. Heterogeneous Multi-task Learning for Human Pose Estimation with Deep Convolutional Neural Network

    Sijin Li, Zhi-Qiang Liu, Antoni B. Chan

    cs.CVcs.LGcs.NEarXiv:1406.3474v12014
  18. Advancing machine learning for MR image reconstruction with an open competition: Overview of the 2019 fastMRI challenge

    Florian Knoll, Tullie Murrell, Anuroop Sriram +8

    eess.IVcs.CVarXiv:2001.02518v12020
  19. Grounding Referring Expressions in Images by Variational Context

    Hanwang Zhang, Yulei Niu, Shih-Fu Chang

    cs.CVarXiv:1712.01892v22017
  20. End-to-End Object Detection with Fully Convolutional Network

    Jianfeng Wang, Lin Song, Zeming Li +3

    cs.CVcs.AIarXiv:2012.03544v32020
  21. Improving DNN Robustness to Adversarial Attacks using Jacobian Regularization

    Daniel Jakubovitz, Raja Giryes

    cs.LGcs.CRcs.CVarXiv:1803.08680v42018
  22. Cluster Alignment with a Teacher for Unsupervised Domain Adaptation

    Zhijie Deng, Yucen Luo, Jun Zhu

    cs.CVarXiv:1903.09980v22019
  23. LR-GAN: Layered Recursive Generative Adversarial Networks for Image Generation

    Jianwei Yang, Anitha Kannan, Dhruv Batra +1

    cs.CVcs.LGarXiv:1703.01560v32017
  24. Perceive, Predict, and Plan: Safe Motion Planning Through Interpretable Semantic Representations

    Abbas Sadat, Sergio Casas, Mengye Ren +3

    cs.ROcs.AIcs.CVarXiv:2008.05930v12020
  25. Detection of Unauthorized IoT Devices Using Machine Learning Techniques

    Yair Meidan, Michael Bohadana, Asaf Shabtai +4

    cs.CRcs.CVarXiv:1709.04647v12017
  26. Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

    Weiyun Wang, Zhe Chen, Wenhai Wang +8

    cs.CLcs.CVarXiv:2411.10442v22024
  27. EndNet: Sparse AutoEncoder Network for Endmember Extraction and Hyperspectral Unmixing

    Savas Ozkan, Berk Kaya, Gozde Bozdagi Akar

    cs.CVarXiv:1708.01894v42017
  28. Articulated Shape Matching Using Laplacian Eigenfunctions and Unsupervised Point Registration

    Diana Mateus, Radu Horaud, David Knossow +2

    cs.CVarXiv:2012.07340v12020
  29. Event-based High Dynamic Range Image and Very High Frame Rate Video Generation using Conditional Generative Adversarial Networks

    S. Mohammad Mostafavi I., Lin Wang, Yo-Sung Ho +1

    cs.CVarXiv:1811.08230v12018
  30. MaskFlownet: Asymmetric Feature Matching with Learnable Occlusion Mask

    Shengyu Zhao, Yilun Sheng, Yue Dong +2

    cs.CVarXiv:2003.10955v22020
  31. Scene Text Recognition from Two-Dimensional Perspective

    Minghui Liao, Jian Zhang, Zhaoyi Wan +5

    cs.CVarXiv:1809.06508v22018
  32. Re-Identification with Consistent Attentive Siamese Networks

    Meng Zheng, Srikrishna Karanam, Ziyan Wu +1

    cs.CVcs.LGarXiv:1811.07487v42018
  33. Generating Diverse Structure for Image Inpainting With Hierarchical VQ-VAE

    Jialun Peng, Dong Liu, Songcen Xu +1

    cs.CVarXiv:2103.10022v12021
  34. Cross-modal Scene Graph Matching for Relationship-aware Image-Text Retrieval

    Sijin Wang, Ruiping Wang, Ziwei Yao +2

    cs.CVarXiv:1910.05134v12019
  35. Continuous-time Intensity Estimation Using Event Cameras

    Cedric Scheerlinck, Nick Barnes, Robert Mahony

    cs.CVarXiv:1811.00386v12018
  36. mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video

    Haiyang Xu, Qinghao Ye, Ming Yan +12

    cs.CVcs.CLcs.MMarXiv:2302.00402v12023
  37. IDNet: Smartphone-based Gait Recognition with Convolutional Neural Networks

    Matteo Gadaleta, Michele Rossi

    cs.CVcs.LGarXiv:1606.03238v32016
  38. Deep Learning for Detecting Building Defects Using Convolutional Neural Networks

    Husein Perez, Joseph H. M. Tah, Amir Mosavi

    cs.CVcs.AIcs.LGarXiv:1908.04392v12019
  39. SparseNeuS: Fast Generalizable Neural Surface Reconstruction from Sparse Views

    Xiaoxiao Long, Cheng Lin, Peng Wang +2

    cs.CVarXiv:2206.05737v22022
  40. Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization

    Zhiyuan Zhao, Bin Wang, Linke Ouyang +3

    cs.CVcs.CLarXiv:2311.16839v22023
  41. Points2Surf: Learning Implicit Surfaces from Point Cloud Patches

    Philipp Erler, Paul Guerrero, Stefan Ohrhallinger +2

    cs.CVarXiv:2007.10453v22020
  42. Adaptive foveated single-pixel imaging with dynamic super-sampling

    David B. Phillips, Ming-Jie Sun, Jonathan M. Taylor +4

    cs.CVphysics.opticsarXiv:1607.08236v12016
  43. Survey on Deep Multi-modal Data Analytics: Collaboration, Rivalry and Fusion

    Yang Wang

    cs.CVarXiv:2006.08159v12020
  44. CTNet: Context-based Tandem Network for Semantic Segmentation

    Zechao Li, Yanpeng Sun, Jinhui Tang

    cs.CVarXiv:2104.09805v12021
  45. Deep Learning with Domain Adaptation for Accelerated Projection-Reconstruction MR

    Yo Seob Han, Jaejun Yoo, Jong Chul Ye

    cs.CVarXiv:1703.01135v22017
  46. Deep Tracking: Seeing Beyond Seeing Using Recurrent Neural Networks

    Peter Ondruska, Ingmar Posner

    cs.LGcs.AIcs.CVarXiv:1602.00991v22016
  47. Automatic Liver Segmentation Using an Adversarial Image-to-Image Network

    Dong Yang, Daguang Xu, S. Kevin Zhou +5

    cs.CVarXiv:1707.08037v12017
  48. StableRep: Synthetic Images from Text-to-Image Models Make Strong Visual Representation Learners

    Yonglong Tian, Lijie Fan, Phillip Isola +2

    cs.CVarXiv:2306.00984v22023
  49. Steganographic Generative Adversarial Networks

    Denis Volkhonskiy, Ivan Nazarov, Evgeny Burnaev

    cs.MMcs.CRcs.CVarXiv:1703.05502v22017
  50. On Pre-Trained Image Features and Synthetic Images for Deep Learning

    Stefan Hinterstoisser, Vincent Lepetit, Paul Wohlhart +1

    cs.CVarXiv:1710.10710v22017
  51. MAP-Net: Multi Attending Path Neural Network for Building Footprint Extraction from Remote Sensed Imagery

    Qing Zhu, Cheng Liao, Han Hu +2

    cs.CVarXiv:1910.12060v22019
  52. MeshDiffusion: Score-based Generative 3D Mesh Modeling

    Zhen Liu, Yao Feng, Michael J. Black +3

    cs.GRcs.AIcs.CVarXiv:2303.08133v22023
  53. Neural Articulated Radiance Field

    Atsuhiro Noguchi, Xiao Sun, Stephen Lin +1

    cs.CVarXiv:2104.03110v22021
  54. Automated Methods for Detection and Classification Pneumonia based on X-Ray Images Using Deep Learning

    Khalid El Asnaoui, Youness Chawki, Ali Idri

    eess.IVcs.CVarXiv:2003.14363v12020
  55. Multi-modal Understanding and Generation for Medical Images and Text via Vision-Language Pre-Training

    Jong Hak Moon, Hyungyung Lee, Woncheol Shin +2

    cs.CVarXiv:2105.11333v32021
  56. Feature-Fused SSD: Fast Detection for Small Objects

    Guimei Cao, Xuemei Xie, Wenzhe Yang +3

    cs.CVarXiv:1709.05054v32017
  57. Fast Point Transformer

    Chunghyun Park, Yoonwoo Jeong, Minsu Cho +1

    cs.CVarXiv:2112.04702v22021
  58. SSAP: Single-Shot Instance Segmentation With Affinity Pyramid

    Naiyu Gao, Yanhu Shan, Yupei Wang +4

    cs.CVarXiv:1909.01616v12019
  59. CurveLane-NAS: Unifying Lane-Sensitive Architecture Search and Adaptive Point Blending

    Hang Xu, Shaoju Wang, Xinyue Cai +3

    cs.CVarXiv:2007.12147v12020
  60. sRGB Real Noise Modeling via Noise-Aware Sampling with Normalizing Flows

    Dongjin Kim, Donggoo Jung, Sungyong Baik +1

    cs.CVarXiv:2608.29038v12026