Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

2,761 to 2,820 of 18,866

  1. EigenPlaces: Training Viewpoint Robust Models for Visual Place Recognition

    Gabriele Berton, Gabriele Trivigno, Barbara Caputo +1

    cs.CVarXiv:2308.10832v12023
  2. Editing Text in the Wild

    Liang Wu, Chengquan Zhang, Jiaming Liu +4

    cs.CVarXiv:1908.03047v12019
  3. Efficient Two-Stage Detection of Human-Object Interactions with a Novel Unary-Pairwise Transformer

    Frederic Z. Zhang, Dylan Campbell, Stephen Gould

    cs.CVcs.AIcs.LGarXiv:2112.01838v22021
  4. Multimodal Optimal Transport-based Co-Attention Transformer with Global Structure Consistency for Survival Prediction

    Yingxue Xu, Hao Chen

    cs.CVarXiv:2306.08330v22023
  5. Crowdsourcing in Computer Vision

    Adriana Kovashka, Olga Russakovsky, Li Fei-Fei +1

    cs.CVcs.HCarXiv:1611.02145v12016
  6. Affective Image Content Analysis: Two Decades Review and New Perspectives

    Sicheng Zhao, Xingxu Yao, Jufeng Yang +5

    cs.CVcs.AIcs.MMarXiv:2106.16125v12021
  7. Matching-CNN Meets KNN: Quasi-Parametric Human Parsing

    Si Liu, Xiaodan Liang, Luoqi Liu +6

    cs.CVarXiv:1504.01220v12015
  8. Breaking Darknet CAPTCHAs with general purpose LLM

    Benjamin Fehrensen, Jens Hubler

    cs.CRcs.CVarXiv:2608.28794v12026
  9. Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection

    Xuechao Zou, Yi Zhou, Kai Li +4

    cs.CVcs.AIarXiv:2609.07670v12026
  10. Towards Universal Representation Learning for Deep Face Recognition

    Yichun Shi, Xiang Yu, Kihyuk Sohn +2

    cs.CVarXiv:2002.11841v12020
  11. Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation

    Igor Pavlovic, Thiemo Wandel, Anton Obukhov +6

    cs.CVcs.LGarXiv:2609.08084v12026
  12. LoCoOp: Few-Shot Out-of-Distribution Detection via Prompt Learning

    Atsuyuki Miyai, Qing Yu, Go Irie +1

    cs.CVarXiv:2306.01293v32023
  13. SC^2-PCR: A Second Order Spatial Compatibility for Efficient and Robust Point Cloud Registration

    Zhi Chen, Kun Sun, Fan Yang +1

    cs.CVarXiv:2203.14453v12022
  14. OadTR: Online Action Detection with Transformers

    Xiang Wang, Shiwei Zhang, Zhiwu Qing +4

    cs.CVarXiv:2106.11149v12021
  15. Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation

    Jaemin Cho, Yushi Hu, Roopal Garg +6

    cs.CVcs.AIcs.CLarXiv:2310.18235v42023
  16. J$\hat{\text{A}}$A-Net: Joint Facial Action Unit Detection and Face Alignment via Adaptive Attention

    Zhiwen Shao, Zhilei Liu, Jianfei Cai +1

    cs.CVarXiv:2003.08834v32020
  17. A survey of advances in vision-based vehicle re-identification

    Sultan Daud Khan, Habib Ullah

    cs.CVcs.AIarXiv:1905.13258v12019
  18. Transformers and Large Language Models for Efficient Intrusion Detection Systems: A Comprehensive Survey

    Hamza Kheddar

    cs.CRcs.AIcs.CLarXiv:2408.07583v22024
  19. RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting

    Hejun Wang, Jinxi Li, Junwei Jiang +4

    cs.CVcs.AIcs.GRarXiv:2609.07414v12026
  20. DuPLO: A DUal view Point deep Learning architecture for time series classificatiOn

    Roberto Interdonato, Dino Ienco, Raffaele Gaetano +1

    cs.CVarXiv:1809.07589v12018
  21. ktrain: A Low-Code Library for Augmented Machine Learning

    Arun S. Maiya

    cs.LGcs.CLcs.CVarXiv:2004.10703v52020
  22. KNN-Diffusion: Image Generation via Large-Scale Retrieval

    Shelly Sheynin, Oron Ashual, Adam Polyak +4

    cs.CVcs.AIcs.CLarXiv:2204.02849v22022
  23. VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

    Sherwin Bahmani, Ivan Skorokhodov, Aliaksandr Siarohin +9

    cs.CVarXiv:2407.12781v32024
  24. Tagger: Deep Unsupervised Perceptual Grouping

    Klaus Greff, Antti Rasmus, Mathias Berglund +3

    cs.CVcs.NEarXiv:1606.06724v22016
  25. Computer aided detection of tuberculosis on chest radiographs: An evaluation of the CAD4TB v6 system

    Keelin Murphy, Shifa Salman Habib, Syed Mohammad Asad Zaidi +10

    eess.IVcs.CVarXiv:1903.03349v22019
  26. GAUDI: A Neural Architect for Immersive 3D Scene Generation

    Miguel Angel Bautista, Pengsheng Guo, Samira Abnar +9

    cs.CVcs.GRcs.LGarXiv:2207.13751v12022
  27. FloorNet: A Unified Framework for Floorplan Reconstruction from 3D Scans

    Chen Liu, Jiaye Wu, Yasutaka Furukawa

    cs.CVarXiv:1804.00090v12018
  28. Attention-driven Graph Clustering Network

    Zhihao Peng, Hui Liu, Yuheng Jia +1

    cs.CVcs.MMarXiv:2108.05499v12021
  29. Latency-Aware Collaborative Perception

    Zixing Lei, Shunli Ren, Yue Hu +2

    cs.CVcs.ROarXiv:2207.08560v42022
  30. ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding

    Chia-Hui Chen, Shih-Ying Yeh, Fu-En Yang +2

    cs.CVarXiv:2609.07941v12026
  31. Improving Image Captioning with Better Use of Captions

    Zhan Shi, Xu Zhou, Xipeng Qiu +1

    cs.CVcs.CLarXiv:2006.11807v12020
  32. Human Pose Estimation using Deep Consensus Voting

    Ita Lifshitz, Ethan Fetaya, Shimon Ullman

    cs.CVcs.LGarXiv:1603.08212v12016
  33. Unsupervised Perceptual Rewards for Imitation Learning

    Pierre Sermanet, Kelvin Xu, Sergey Levine

    cs.CVcs.ROarXiv:1612.06699v32016
  34. Foundational Models Defining a New Era in Vision: A Survey and Outlook

    Muhammad Awais, Muzammal Naseer, Salman Khan +5

    cs.CVcs.AIarXiv:2307.13721v12023
  35. Reason2Drive: Towards Interpretable and Chain-based Reasoning for Autonomous Driving

    Ming Nie, Renyuan Peng, Chunwei Wang +4

    cs.CVarXiv:2312.03661v32023
  36. MonoRUn: Monocular 3D Object Detection by Reconstruction and Uncertainty Propagation

    Hansheng Chen, Yuyao Huang, Wei Tian +2

    cs.CVarXiv:2103.12605v22021
  37. Cars Can't Fly up in the Sky: Improving Urban-Scene Segmentation via Height-driven Attention Networks

    Sungha Choi, Joanne T. Kim, Jaegul Choo

    cs.CVarXiv:2003.05128v32020
  38. Graph Structure of Neural Networks

    Jiaxuan You, Jure Leskovec, Kaiming He +1

    cs.LGcs.CVcs.SIarXiv:2007.06559v22020
  39. Localizing Object-level Shape Variations with Text-to-Image Diffusion Models

    Or Patashnik, Daniel Garibi, Idan Azuri +2

    cs.CVcs.GRcs.LGarXiv:2303.11306v22023
  40. HeadGAN: One-shot Neural Head Synthesis and Editing

    Michail Christos Doukas, Stefanos Zafeiriou, Viktoriia Sharmanska

    cs.CVarXiv:2012.08261v32020
  41. Contrastive Knowledge Distillation for Anomaly Detection in Multi-Illumination/Focus Display Images

    Jihyun Lee, Hangil Park, Yongmin Seo +4

    cs.CVcs.LGarXiv:2609.05520v12026
  42. CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs

    Nhat-Tan Bui, Varshini Elangovan, Arun Reddy Anugu +7

    cs.CVcs.LGarXiv:2609.08345v12026
  43. AlphaPilot: Autonomous Drone Racing

    Philipp Foehn, Dario Brescianini, Elia Kaufmann +4

    cs.ROcs.CVeess.SYarXiv:2005.12813v22020
  44. Do Depressive Facial Patterns Transfer Across Cultures and Contexts? Evidence from a German RCT and E-DAIC

    Misha Sadeghi, Robert Richer, Lydia Helene Rupp +7

    cs.CVcs.HCarXiv:2609.05543v12026
  45. Exploiting Recurrent Neural Networks and Leap Motion Controller for Sign Language and Semaphoric Gesture Recognition

    Danilo Avola, Marco Bernardi, Luigi Cinque +2

    cs.CVarXiv:1803.10435v12018
  46. Universal Humanoid Motion Representations for Physics-Based Control

    Zhengyi Luo, Jinkun Cao, Josh Merel +4

    cs.CVcs.GRcs.ROarXiv:2310.04582v22023
  47. Situation Awareness for Intelligent Data Distribution in Connected Vehicles

    Falk Dettinger, Akshay Narla, Michael Weyrich

    cs.ROcs.AIcs.CVarXiv:2609.05521v12026
  48. Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through Estimation

    Zechun Liu, Kwang-Ting Cheng, Dong Huang +2

    cs.CVcs.AIcs.LGarXiv:2111.14826v22021
  49. Representation learning of human cortical folding to reveal long lasting neurodevelopmental signatures

    Julien Laval, Robin Guiavarch, Antoine Dufournet +24

    q-bio.QMcs.CVcs.LGarXiv:2609.05438v12026
  50. Coarse-to-Fine Vision-Language Pre-training with Fusion in the Backbone

    Zi-Yi Dou, Aishwarya Kamath, Zhe Gan +9

    cs.CVcs.CLcs.LGarXiv:2206.07643v22022
  51. Efficient Learning on Point Clouds with Basis Point Sets

    Sergey Prokudin, Christoph Lassner, Javier Romero

    cs.CVarXiv:1908.09186v12019
  52. SelfOcc: Self-Supervised Vision-Based 3D Occupancy Prediction

    Yuanhui Huang, Wenzhao Zheng, Borui Zhang +2

    cs.CVcs.AIcs.LGarXiv:2311.12754v22023
  53. FCA: Learning a 3D Full-coverage Vehicle Camouflage for Multi-view Physical Adversarial Attack

    Donghua Wang, Tingsong Jiang, Jialiang Sun +5

    cs.CVcs.AIarXiv:2109.07193v32021
  54. SurgicalSAM: Efficient Class Promptable Surgical Instrument Segmentation

    Wenxi Yue, Jing Zhang, Kun Hu +3

    cs.CVcs.AIcs.ROarXiv:2308.08746v22023
  55. Perceptual Quality Assessment of Omnidirectional Images

    Huiyu Duan, Guangtao Zhai, Xiongkuo Min +3

    cs.CVarXiv:2207.02674v12022
  56. Streamlined Dense Video Captioning

    Jonghwan Mun, Linjie Yang, Zhou Ren +2

    cs.CVarXiv:1904.03870v12019
  57. A Deep Pyramid Deformable Part Model for Face Detection

    Rajeev Ranjan, Vishal M. Patel, Rama Chellappa

    cs.CVarXiv:1508.04389v12015
  58. Video Cloze Procedure for Self-Supervised Spatio-Temporal Learning

    Dezhao Luo, Chang Liu, Yu Zhou +4

    cs.CVarXiv:2001.00294v12020
  59. rPPG-Toolbox: Deep Remote PPG Toolbox

    Xin Liu, Girish Narayanswamy, Akshay Paruchuri +7

    cs.CVarXiv:2210.00716v32022
  60. On Learning 3D Face Morphable Model from In-the-wild Images

    Luan Tran, Xiaoming Liu

    cs.CVarXiv:1808.09560v22018