Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

9,601 to 9,660 of 18,776

  1. Illuminating Pedestrians via Simultaneous Detection & Segmentation

    Garrick Brazil, Xi Yin, Xiaoming Liu

    cs.CVarXiv:1706.08564v12017
  2. Visual Question Answering: Datasets, Algorithms, and Future Challenges

    Kushal Kafle, Christopher Kanan

    cs.CVcs.AIcs.CLarXiv:1610.01465v42016
  3. ET-Net: A Generic Edge-aTtention Guidance Network for Medical Image Segmentation

    Zhijie Zhang, Huazhu Fu, Hang Dai +3

    cs.CVarXiv:1907.10936v12019
    Summaries:한국어
  4. Adversarial Diversity and Hard Positive Generation

    Andras Rozsa, Ethan M. Rudd, Terrance E. Boult

    cs.CVarXiv:1605.01775v22016
  5. TransNet V2: An effective deep network architecture for fast shot transition detection

    Tomáš Souček, Jakub Lokoč

    cs.CVarXiv:2008.04838v12020
  6. Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

    Jianwei Yang, Hao Zhang, Feng Li +3

    cs.CVcs.AIcs.CLarXiv:2310.11441v22023
  7. MemSeg: A semi-supervised method for image surface defect detection using differences and commonalities

    Minghui Yang, Peng Wu, Jing Liu +1

    cs.CVarXiv:2205.00908v12022
  8. X2CT-GAN: Reconstructing CT from Biplanar X-Rays with Generative Adversarial Networks

    Xingde Ying, Heng Guo, Kai Ma +3

    eess.IVcs.CVarXiv:1905.06902v12019
  9. Structured Feature Learning for Pose Estimation

    Xiao Chu, Wanli Ouyang, Hongsheng Li +1

    cs.CVarXiv:1603.09065v12016
  10. Affect Analysis in-the-wild: Valence-Arousal, Expressions, Action Units and a Unified Framework

    Dimitrios Kollias, Stefanos Zafeiriou

    cs.CVcs.AIcs.LGarXiv:2103.15792v12021
  11. Gotta Go Fast When Generating Data with Score-Based Models

    Alexia Jolicoeur-Martineau, Ke Li, Rémi Piché-Taillefer +2

    cs.LGcs.CVmath.OCarXiv:2105.14080v12021
  12. Self-Supervised Video Hashing with Hierarchical Binary Auto-encoder

    Jingkuan Song, Hanwang Zhang, Xiangpeng Li +3

    cs.CVarXiv:1802.02305v12018
  13. Do Datasets Have Politics? Disciplinary Values in Computer Vision Dataset Development

    Morgan Klaus Scheuerman, Emily Denton, Alex Hanna

    cs.CVcs.HCarXiv:2108.04308v22021
  14. A Hybrid Deep Learning Architecture for Privacy-Preserving Mobile Analytics

    Seyed Ali Osia, Ali Shahin Shamsabadi, Sina Sajadmanesh +5

    cs.LGcs.CVarXiv:1703.02952v72017
  15. Kernel Methods on Riemannian Manifolds with Gaussian RBF Kernels

    Sadeep Jayasumana, Richard Hartley, Mathieu Salzmann +2

    cs.CVarXiv:1412.0265v22014
  16. Learning by Aligning: Visible-Infrared Person Re-identification using Cross-Modal Correspondences

    Hyunjong Park, Sanghoon Lee, Junghyup Lee +1

    cs.CVarXiv:2108.07422v12021
  17. FixBi: Bridging Domain Spaces for Unsupervised Domain Adaptation

    Jaemin Na, Heechul Jung, Hyung Jin Chang +1

    cs.CVarXiv:2011.09230v22020
  18. Two-Stream Network for Sign Language Recognition and Translation

    Yutong Chen, Ronglai Zuo, Fangyun Wei +3

    cs.CVarXiv:2211.01367v22022
  19. See Better Before Looking Closer: Weakly Supervised Data Augmentation Network for Fine-Grained Visual Classification

    Tao Hu, Honggang Qi, Qingming Huang +1

    cs.CVarXiv:1901.09891v22019
  20. FourLLIE: Boosting Low-Light Image Enhancement by Fourier Frequency Information

    Chenxi Wang, Hongjun Wu, Zhi Jin

    cs.CVeess.IVarXiv:2308.03033v12023
  21. Rethinking on Multi-Stage Networks for Human Pose Estimation

    Wenbo Li, Zhicheng Wang, Binyi Yin +7

    cs.CVarXiv:1901.00148v42019
  22. RIO: 3D Object Instance Re-Localization in Changing Indoor Environments

    Johanna Wald, Armen Avetisyan, Nassir Navab +2

    cs.CVarXiv:1908.06109v12019
  23. ACRONYM: A Large-Scale Grasp Dataset Based on Simulation

    Clemens Eppner, Arsalan Mousavian, Dieter Fox

    cs.ROcs.CVarXiv:2011.09584v12020
  24. NumBench: Diagnosing Counting Failures in Text-to-Image Models

    Sandeep Wadhwa, Mayank Vatsa, Richa Singh +2

    cs.CVcs.DBarXiv:2608.28206v12026
  25. Crowd counting via scale-adaptive convolutional neural network

    Lu Zhang, Miaojing Shi, Qiaobo Chen

    cs.CVarXiv:1711.04433v42017
  26. The Sound of Motions

    Hang Zhao, Chuang Gan, Wei-Chiu Ma +1

    cs.CVcs.SDeess.ASarXiv:1904.05979v12019
  27. Fast End-to-End Trainable Guided Filter

    Huikai Wu, Shuai Zheng, Junge Zhang +1

    cs.CVarXiv:1803.05619v22018
  28. Universal Instance Perception as Object Discovery and Retrieval

    Bin Yan, Yi Jiang, Jiannan Wu +4

    cs.CVarXiv:2303.06674v22023
  29. Self-Chained Image-Language Model for Video Localization and Question Answering

    Shoubin Yu, Jaemin Cho, Prateek Yadav +1

    cs.CVcs.AIcs.CLarXiv:2305.06988v22023
  30. 3D Reconstruction with Spatial Memory

    Hengyi Wang, Lourdes Agapito

    cs.CVarXiv:2408.16061v12024
  31. End-to-End Human Object Interaction Detection with HOI Transformer

    Cheng Zou, Bohan Wang, Yue Hu +8

    cs.CVarXiv:2103.04503v12021
  32. Pinwheel-shaped Convolution and Scale-based Dynamic Loss for Infrared Small Target Detection

    Jiangnan Yang, Shuangli Liu, Jingjun Wu +3

    cs.CVarXiv:2412.16986v12024
  33. DensityKV: Density-Guided KV Cache Compression for Long Video Generation

    Wenqu Zhao, Xuemin Chi, Xin Zhang +6

    cs.CVarXiv:2608.27922v12026
  34. Equivariant Multi-Modality Image Fusion

    Zixiang Zhao, Haowen Bai, Jiangshe Zhang +6

    cs.CVarXiv:2305.11443v22023
  35. Towards Visually Explaining Variational Autoencoders

    Wenqian Liu, Runze Li, Meng Zheng +5

    cs.CVcs.LGarXiv:1911.07389v72019
  36. Federated Learning for Medical Applications: A Taxonomy, Current Trends, Challenges, and Future Research Directions

    Ashish Rauniyar, Desta Haileselassie Hagos, Debesh Jha +4

    cs.LGcs.CRcs.CVarXiv:2208.03392v52022
  37. Efficient and Accurate Approximations of Nonlinear Convolutional Networks

    Xiangyu Zhang, Jianhua Zou, Xiang Ming +2

    cs.CVarXiv:1411.4229v12014
  38. HyperDreamBooth: HyperNetworks for Fast Personalization of Text-to-Image Models

    Nataniel Ruiz, Yuanzhen Li, Varun Jampani +6

    cs.CVcs.AIcs.GRarXiv:2307.06949v22023
  39. Geometric Feature-Based Facial Expression Recognition in Image Sequences Using Multi-Class AdaBoost and Support Vector Machines

    Deepak Ghimire, Joonwhoan Lee

    cs.CVarXiv:1604.03225v12016
  40. Climate Physics Dynamic Matching

    Gurjeet Sangra Singh, Frantzeska Lavda, Alexandros Kalousis

    stat.APcs.CVcs.LOarXiv:2608.26907v12026
  41. ICDAR2019 Robust Reading Challenge on Arbitrary-Shaped Text (RRC-ArT)

    Chee-Kheng Chng, Yuliang Liu, Yipeng Sun +11

    cs.CVarXiv:1909.07145v12019
  42. Self-supervised Learning with Geometric Constraints in Monocular Video: Connecting Flow, Depth, and Camera

    Yuhua Chen, Cordelia Schmid, Cristian Sminchisescu

    cs.CVarXiv:1907.05820v22019
  43. MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

    Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier +29

    cs.CVcs.CLcs.LGarXiv:2403.09611v42024
  44. AutoAssign: Differentiable Label Assignment for Dense Object Detection

    Benjin Zhu, Jianfeng Wang, Zhengkai Jiang +4

    cs.CVarXiv:2007.03496v32020
  45. A Simple Framework for Open-Vocabulary Segmentation and Detection

    Hao Zhang, Feng Li, Xueyan Zou +5

    cs.CVarXiv:2303.08131v32023
  46. Dual-Stream Semantic Guidance with Prototype Anchor Calibration for Source-Fully-Free Adaptation of Vision-Language Models

    Weiwei Xiang, Shun Peng, Guangyi Xiao +2

    cs.CVarXiv:2608.28145v12026
  47. DeepfakeBench: A Comprehensive Benchmark of Deepfake Detection

    Zhiyuan Yan, Yong Zhang, Xinhang Yuan +2

    cs.CVarXiv:2307.01426v22023
  48. PersFormer: 3D Lane Detection via Perspective Transformer and the OpenLane Benchmark

    Li Chen, Chonghao Sima, Yang Li +8

    cs.CVarXiv:2203.11089v32022
  49. Gait Recognition in the Wild with Dense 3D Representations and A Benchmark

    Jinkai Zheng, Xinchen Liu, Wu Liu +3

    cs.CVarXiv:2204.02569v12022
  50. Web-Scale Training for Face Identification

    Yaniv Taigman, Ming Yang, Marc'Aurelio Ranzato +1

    cs.CVarXiv:1406.5266v22014
  51. Physically-Based Rendering for Indoor Scene Understanding Using Convolutional Neural Networks

    Yinda Zhang, Shuran Song, Ersin Yumer +4

    cs.CVarXiv:1612.07429v32016
  52. LucidDreamer: Domain-free Generation of 3D Gaussian Splatting Scenes

    Jaeyoung Chung, Suyoung Lee, Hyeongjin Nam +2

    cs.CVarXiv:2311.13384v22023
  53. Missing MRI Pulse Sequence Synthesis using Multi-Modal Generative Adversarial Network

    Anmol Sharma, Ghassan Hamarneh

    eess.IVcs.AIcs.CVarXiv:1904.12200v32019
  54. Universal Litmus Patterns: Revealing Backdoor Attacks in CNNs

    Soheil Kolouri, Aniruddha Saha, Hamed Pirsiavash +1

    cs.CVarXiv:1906.10842v22019
  55. Geometry-Consistent Generative Adversarial Networks for One-Sided Unsupervised Domain Mapping

    Huan Fu, Mingming Gong, Chaohui Wang +3

    cs.CVarXiv:1809.05852v22018
  56. D3S -- A Discriminative Single Shot Segmentation Tracker

    Alan Lukežič, Jiří Matas, Matej Kristan

    cs.CVarXiv:1911.08862v22019
  57. Dual Residual Networks Leveraging the Potential of Paired Operations for Image Restoration

    Xing Liu, Masanori Suganuma, Zhun Sun +1

    cs.CVarXiv:1903.08817v22019
  58. A Unified Objective for Novel Class Discovery

    Enrico Fini, Enver Sangineto, Stéphane Lathuilière +3

    cs.CVcs.LGarXiv:2108.08536v42021
  59. Learning Video Representations from Large Language Models

    Yue Zhao, Ishan Misra, Philipp Krähenbühl +1

    cs.CVarXiv:2212.04501v12022
  60. A Comprehensive Review of Computer-aided Whole-slide Image Analysis: from Datasets to Feature Extraction, Segmentation, Classification, and Detection Approaches

    Chen Li, Xintong Li, Md Rahaman +8

    cs.CVcs.AIarXiv:2102.10553v12021