Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

2,041 to 2,100 of 18,830

  1. A Two-Stage Method for Text Line Detection in Historical Documents

    Tobias Grüning, Gundram Leifert, Tobias Strauß +2

    cs.CVarXiv:1802.03345v22018
  2. Live Face De-Identification in Video

    Oran Gafni, Lior Wolf, Yaniv Taigman

    cs.LGcs.CVcs.GRarXiv:1911.08348v12019
  3. Continuous 3D Label Stereo Matching using Local Expansion Moves

    Tatsunori Taniai, Yasuyuki Matsushita, Yoichi Sato +1

    cs.CVarXiv:1603.08328v32016
  4. EFANNA : An Extremely Fast Approximate Nearest Neighbor Search Algorithm Based on kNN Graph

    Cong Fu, Deng Cai

    cs.CVarXiv:1609.07228v32016
  5. Learning Multi-Granular Hypergraphs for Video-Based Person Re-Identification

    Yichao Yan, Jie Qin1, Jiaxin Chen +4

    cs.CVarXiv:2104.14913v12021
  6. TS2C: Tight Box Mining with Surrounding Segmentation Context for Weakly Supervised Object Detection

    Yunchao Wei, Zhiqiang Shen, Bowen Cheng +4

    cs.CVarXiv:1807.04897v12018
  7. Learning to Listen: Modeling Non-Deterministic Dyadic Facial Motion

    Evonne Ng, Hanbyul Joo, Liwen Hu +4

    cs.CVarXiv:2204.08451v12022
  8. IAN: The Individual Aggregation Network for Person Search

    Jimin Xiao, Yanchun Xie, Tammam Tillo +3

    cs.CVarXiv:1705.05552v12017
  9. Deep Parametric Indoor Lighting Estimation

    Marc-André Gardner, Yannick Hold-Geoffroy, Kalyan Sunkavalli +2

    cs.CVarXiv:1910.08812v12019
  10. Progressive Face Super-Resolution via Attention to Facial Landmark

    Deokyun Kim, Minseon Kim, Gihyun Kwon +1

    cs.CVarXiv:1908.08239v12019
  11. Zoom-Net: Mining Deep Feature Interactions for Visual Relationship Recognition

    Guojun Yin, Lu Sheng, Bin Liu +4

    cs.CVarXiv:1807.04979v12018
  12. Unsupervised Part-Based Disentangling of Object Shape and Appearance

    Dominik Lorenz, Leonard Bereska, Timo Milbich +1

    cs.CVarXiv:1903.06946v32019
  13. Learning Hierarchical Cross-Modal Association for Co-Speech Gesture Generation

    Xian Liu, Qianyi Wu, Hang Zhou +7

    cs.CVarXiv:2203.13161v12022
  14. Photo-Sketching: Inferring Contour Drawings from Images

    Mengtian Li, Zhe Lin, Radomir Mech +2

    cs.CVarXiv:1901.00542v12019
  15. And the Bit Goes Down: Revisiting the Quantization of Neural Networks

    Pierre Stock, Armand Joulin, Rémi Gribonval +2

    cs.CVarXiv:1907.05686v52019
  16. DG-Font: Deformable Generative Networks for Unsupervised Font Generation

    Yangchen Xie, Xinyuan Chen, Li Sun +1

    cs.CVarXiv:2104.03064v22021
  17. Attribute Recognition by Joint Recurrent Learning of Context and Correlation

    Jingya Wang, Xiatian Zhu, Shaogang Gong +1

    cs.CVarXiv:1709.08553v12017
  18. Discriminative Feature Learning for Unsupervised Video Summarization

    Yunjae Jung, Donghyeon Cho, Dahun Kim +2

    cs.CVarXiv:1811.09791v12018
  19. Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens

    Zhangqi Jiang, Junkai Chen, Beier Zhu +3

    cs.CVarXiv:2411.16724v32024
  20. Spatially Resolved Gene Expression Prediction from H&E Histology Images via Bi-modal Contrastive Learning

    Ronald Xie, Kuan Pang, Sai W. Chung +4

    cs.CVcs.AIarXiv:2306.01859v22023
  21. Dinomaly: The Less Is More Philosophy in Multi-Class Unsupervised Anomaly Detection

    Jia Guo, Shuai Lu, Weihang Zhang +3

    cs.CVarXiv:2405.14325v52024
  22. A Comparison of Nature Inspired Algorithms for Multi-threshold Image Segmentation

    Valentín Osuna-Enciso, Erik Cuevas, Humberto Sossa

    cs.CVcs.NEarXiv:1405.7406v12014
  23. Correlation-aware Adversarial Domain Adaptation and Generalization

    Mohammad Mahfujur Rahman, Clinton Fookes, Mahsa Baktashmotlagh +1

    cs.CVcs.LGarXiv:1911.12983v12019
  24. Dreaming More Data: Class-dependent Distributions over Diffeomorphisms for Learned Data Augmentation

    Søren Hauberg, Oren Freifeld, Anders Boesen Lindbo Larsen +2

    cs.CVarXiv:1510.02795v22015
  25. Deep Learning in Cardiology

    Paschalis Bizopoulos, Dimitrios Koutsouris

    cs.CVcs.AIcs.LGarXiv:1902.11122v52019
  26. Scalable Vision Transformers with Hierarchical Pooling

    Zizheng Pan, Bohan Zhuang, Jing Liu +2

    cs.CVarXiv:2103.10619v22021
  27. SweetDreamer: Aligning Geometric Priors in 2D Diffusion for Consistent Text-to-3D

    Weiyu Li, Rui Chen, Xuelin Chen +1

    cs.CVarXiv:2310.02596v22023
  28. A cross Transformer for image denoising

    Chunwei Tian, Menghua Zheng, Wangmeng Zuo +3

    eess.IVcs.CVcs.LGarXiv:2310.10408v12023
  29. From One Hand to Multiple Hands: Imitation Learning for Dexterous Manipulation from Single-Camera Teleoperation

    Yuzhe Qin, Hao Su, Xiaolong Wang

    cs.ROcs.CVcs.LGarXiv:2204.12490v22022
  30. Deep Model Reassembly

    Xingyi Yang, Daquan Zhou, Songhua Liu +2

    cs.CVcs.AIarXiv:2210.17409v22022
  31. Building Fast and Compact Convolutional Neural Networks for Offline Handwritten Chinese Character Recognition

    Xuefeng Xiao, Lianwen Jin, Yafeng Yang +3

    cs.CVarXiv:1702.07975v12017
  32. Mono-Camera 3D Multi-Object Tracking Using Deep Learning Detections and PMBM Filtering

    Samuel Scheidegger, Joachim Benjaminsson, Emil Rosenberg +2

    cs.CVeess.SParXiv:1802.09975v12018
  33. Joint Semantic Segmentation and Depth Estimation with Deep Convolutional Networks

    Arsalan Mousavian, Hamed Pirsiavash, Jana Kosecka

    cs.CVarXiv:1604.07480v32016
  34. DCReg: Decoupled Characterization for Efficient Degenerate LiDAR Registration

    Xiangcheng Hu, Xieyuanli Chen, Mingkai Jia +3

    cs.ROcs.CVarXiv:2509.06285v22025
  35. Learning Meta Face Recognition in Unseen Domains

    Jianzhu Guo, Xiangyu Zhu, Chenxu Zhao +3

    cs.CVarXiv:2003.07733v22020
  36. NuClick: A Deep Learning Framework for Interactive Segmentation of Microscopy Images

    Navid Alemi Koohbanani, Mostafa Jahanifar, Neda Zamani Tajadin +1

    cs.CVstat.AParXiv:2005.14511v22020
  37. AENet: Learning Deep Audio Features for Video Analysis

    Naoya Takahashi, Michael Gygli, Luc Van Gool

    cs.MMcs.CVcs.SDarXiv:1701.00599v22017
  38. Control Copy-Paste: Controllable Diffusion-Based Augmentation Method for Remote Sensing Few-Shot Object Detection

    Yanxing Liu, Jiancheng Pan, Bingchen Zhang

    eess.IVcs.CVarXiv:2507.21816v12025
  39. QDTrack: Quasi-Dense Similarity Learning for Appearance-Only Multiple Object Tracking

    Tobias Fischer, Thomas E. Huang, Jiangmiao Pang +4

    cs.CVarXiv:2210.06984v22022
  40. 3D Anisotropic Hybrid Network: Transferring Convolutional Features from 2D Images to 3D Anisotropic Volumes

    Siqi Liu, Daguang Xu, S. Kevin Zhou +7

    cs.CVarXiv:1711.08580v22017
  41. Concurrent Segmentation and Localization for Tracking of Surgical Instruments

    Iro Laina, Nicola Rieke, Christian Rupprecht +4

    cs.CVarXiv:1703.10701v22017
  42. Enhancing Pseudo Label Quality for Semi-Supervised Domain-Generalized Medical Image Segmentation

    Huifeng Yao, Xiaowei Hu, Xiaomeng Li

    cs.CVcs.AIarXiv:2201.08657v22022
  43. Diverse Instance Generation via Diffusion Models for Enhanced Few-Shot Object Detection in Remote Sensing Images

    Yanxing Liu, Jiancheng Pan, Jianwei Yang +3

    eess.IVcs.CVarXiv:2511.18031v12025
  44. Comparison of Time-Frequency Representations for Environmental Sound Classification using Convolutional Neural Networks

    M. Huzaifah

    cs.CVarXiv:1706.07156v12017
  45. Learning Non-target Knowledge for Few-shot Semantic Segmentation

    Yuanwei Liu, Nian Liu, Qinglong Cao +3

    cs.CVarXiv:2205.04903v12022
  46. Hierarchical Diffusion Policy for Kinematics-Aware Multi-Task Robotic Manipulation

    Xiao Ma, Sumit Patidar, Iain Haughton +1

    cs.ROcs.AIcs.CVarXiv:2403.03890v12024
  47. Color-based Segmentation of Sky/Cloud Images From Ground-based Cameras

    Soumyabrata Dev, Yee Hui Lee, Stefan Winkler

    cs.CVarXiv:1606.03669v12016
  48. UrbanLoco: A Full Sensor Suite Dataset for Mapping and Localization in Urban Scenes

    Weisong Wen, Yiyang Zhou, Guohao Zhang +5

    cs.ROcs.CVarXiv:1912.09513v22019
  49. Self Pre-training with Masked Autoencoders for Medical Image Classification and Segmentation

    Lei Zhou, Huidong Liu, Joseph Bae +3

    eess.IVcs.CVcs.LGarXiv:2203.05573v22022
  50. Attention for Fine-Grained Categorization

    Pierre Sermanet, Andrea Frome, Esteban Real

    cs.CVcs.LGcs.NEarXiv:1412.7054v32014
  51. Towards Principled Disentanglement for Domain Generalization

    Hanlin Zhang, Yi-Fan Zhang, Weiyang Liu +3

    cs.LGcs.CVarXiv:2111.13839v42021
  52. Learning like a Child: Fast Novel Visual Concept Learning from Sentence Descriptions of Images

    Junhua Mao, Wei Xu, Yi Yang +3

    cs.CVcs.CLcs.LGarXiv:1504.06692v22015
  53. CADP: A Novel Dataset for CCTV Traffic Camera based Accident Analysis

    Ankit Shah, Jean Baptiste Lamare, Tuan Nguyen Anh +1

    cs.CVcs.MMarXiv:1809.05782v22018
  54. StyleGAN2 Distillation for Feed-forward Image Manipulation

    Yuri Viazovetskyi, Vladimir Ivashkin, Evgeny Kashin

    cs.CVarXiv:2003.03581v22020
  55. DISC: Deep Image Saliency Computing via Progressive Representation Learning

    Tianshui Chen, Liang Lin, Lingbo Liu +2

    cs.CVarXiv:1511.04192v22015
  56. On Bringing Robots Home

    Nur Muhammad Mahi Shafiullah, Anant Rai, Haritheja Etukuru +4

    cs.ROcs.AIcs.CVarXiv:2311.16098v12023
  57. Incremental learning for the detection and classification of GAN-generated images

    Francesco Marra, Cristiano Saltori, Giulia Boato +1

    cs.CVarXiv:1910.01568v22019
  58. segDeepM: Exploiting Segmentation and Context in Deep Neural Networks for Object Detection

    Yukun Zhu, Raquel Urtasun, Ruslan Salakhutdinov +1

    cs.CVarXiv:1502.04275v12015
  59. Comparing SNNs and RNNs on Neuromorphic Vision Datasets: Similarities and Differences

    Weihua He, YuJie Wu, Lei Deng +6

    cs.CVcs.NEeess.IVarXiv:2005.02183v12020
  60. Temporal Action Segmentation: An Analysis of Modern Techniques

    Guodong Ding, Fadime Sener, Angela Yao

    cs.CVarXiv:2210.10352v52022