Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
2,761 to 2,820 of 18,866
EigenPlaces: Training Viewpoint Robust Models for Visual Place Recognition
Gabriele Berton, Gabriele Trivigno, Barbara Caputo +1
cs.CVarXiv:2308.10832v12023Editing Text in the Wild
Liang Wu, Chengquan Zhang, Jiaming Liu +4
cs.CVarXiv:1908.03047v12019Efficient Two-Stage Detection of Human-Object Interactions with a Novel Unary-Pairwise Transformer
Frederic Z. Zhang, Dylan Campbell, Stephen Gould
cs.CVcs.AIcs.LGarXiv:2112.01838v22021Multimodal Optimal Transport-based Co-Attention Transformer with Global Structure Consistency for Survival Prediction
Yingxue Xu, Hao Chen
cs.CVarXiv:2306.08330v22023Crowdsourcing in Computer Vision
Adriana Kovashka, Olga Russakovsky, Li Fei-Fei +1
cs.CVcs.HCarXiv:1611.02145v12016Affective Image Content Analysis: Two Decades Review and New Perspectives
Sicheng Zhao, Xingxu Yao, Jufeng Yang +5
cs.CVcs.AIcs.MMarXiv:2106.16125v12021Matching-CNN Meets KNN: Quasi-Parametric Human Parsing
Si Liu, Xiaodan Liang, Luoqi Liu +6
cs.CVarXiv:1504.01220v12015Breaking Darknet CAPTCHAs with general purpose LLM
Benjamin Fehrensen, Jens Hubler
cs.CRcs.CVarXiv:2608.28794v12026Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection
Xuechao Zou, Yi Zhou, Kai Li +4
cs.CVcs.AIarXiv:2609.07670v12026Towards Universal Representation Learning for Deep Face Recognition
Yichun Shi, Xiang Yu, Kihyuk Sohn +2
cs.CVarXiv:2002.11841v12020Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation
Igor Pavlovic, Thiemo Wandel, Anton Obukhov +6
cs.CVcs.LGarXiv:2609.08084v12026LoCoOp: Few-Shot Out-of-Distribution Detection via Prompt Learning
Atsuyuki Miyai, Qing Yu, Go Irie +1
cs.CVarXiv:2306.01293v32023SC^2-PCR: A Second Order Spatial Compatibility for Efficient and Robust Point Cloud Registration
Zhi Chen, Kun Sun, Fan Yang +1
cs.CVarXiv:2203.14453v12022OadTR: Online Action Detection with Transformers
Xiang Wang, Shiwei Zhang, Zhiwu Qing +4
cs.CVarXiv:2106.11149v12021Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation
Jaemin Cho, Yushi Hu, Roopal Garg +6
cs.CVcs.AIcs.CLarXiv:2310.18235v42023J$\hat{\text{A}}$A-Net: Joint Facial Action Unit Detection and Face Alignment via Adaptive Attention
Zhiwen Shao, Zhilei Liu, Jianfei Cai +1
cs.CVarXiv:2003.08834v32020A survey of advances in vision-based vehicle re-identification
Sultan Daud Khan, Habib Ullah
cs.CVcs.AIarXiv:1905.13258v12019Transformers and Large Language Models for Efficient Intrusion Detection Systems: A Comprehensive Survey
Hamza Kheddar
cs.CRcs.AIcs.CLarXiv:2408.07583v22024RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting
Hejun Wang, Jinxi Li, Junwei Jiang +4
cs.CVcs.AIcs.GRarXiv:2609.07414v12026DuPLO: A DUal view Point deep Learning architecture for time series classificatiOn
Roberto Interdonato, Dino Ienco, Raffaele Gaetano +1
cs.CVarXiv:1809.07589v12018ktrain: A Low-Code Library for Augmented Machine Learning
Arun S. Maiya
cs.LGcs.CLcs.CVarXiv:2004.10703v52020KNN-Diffusion: Image Generation via Large-Scale Retrieval
Shelly Sheynin, Oron Ashual, Adam Polyak +4
cs.CVcs.AIcs.CLarXiv:2204.02849v22022VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control
Sherwin Bahmani, Ivan Skorokhodov, Aliaksandr Siarohin +9
cs.CVarXiv:2407.12781v32024Tagger: Deep Unsupervised Perceptual Grouping
Klaus Greff, Antti Rasmus, Mathias Berglund +3
cs.CVcs.NEarXiv:1606.06724v22016Computer aided detection of tuberculosis on chest radiographs: An evaluation of the CAD4TB v6 system
Keelin Murphy, Shifa Salman Habib, Syed Mohammad Asad Zaidi +10
eess.IVcs.CVarXiv:1903.03349v22019GAUDI: A Neural Architect for Immersive 3D Scene Generation
Miguel Angel Bautista, Pengsheng Guo, Samira Abnar +9
cs.CVcs.GRcs.LGarXiv:2207.13751v12022FloorNet: A Unified Framework for Floorplan Reconstruction from 3D Scans
Chen Liu, Jiaye Wu, Yasutaka Furukawa
cs.CVarXiv:1804.00090v12018Attention-driven Graph Clustering Network
Zhihao Peng, Hui Liu, Yuheng Jia +1
cs.CVcs.MMarXiv:2108.05499v12021Latency-Aware Collaborative Perception
Zixing Lei, Shunli Ren, Yue Hu +2
cs.CVcs.ROarXiv:2207.08560v42022ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding
Chia-Hui Chen, Shih-Ying Yeh, Fu-En Yang +2
cs.CVarXiv:2609.07941v12026Improving Image Captioning with Better Use of Captions
Zhan Shi, Xu Zhou, Xipeng Qiu +1
cs.CVcs.CLarXiv:2006.11807v12020Human Pose Estimation using Deep Consensus Voting
Ita Lifshitz, Ethan Fetaya, Shimon Ullman
cs.CVcs.LGarXiv:1603.08212v12016Unsupervised Perceptual Rewards for Imitation Learning
Pierre Sermanet, Kelvin Xu, Sergey Levine
cs.CVcs.ROarXiv:1612.06699v32016Foundational Models Defining a New Era in Vision: A Survey and Outlook
Muhammad Awais, Muzammal Naseer, Salman Khan +5
cs.CVcs.AIarXiv:2307.13721v12023Reason2Drive: Towards Interpretable and Chain-based Reasoning for Autonomous Driving
Ming Nie, Renyuan Peng, Chunwei Wang +4
cs.CVarXiv:2312.03661v32023MonoRUn: Monocular 3D Object Detection by Reconstruction and Uncertainty Propagation
Hansheng Chen, Yuyao Huang, Wei Tian +2
cs.CVarXiv:2103.12605v22021Cars Can't Fly up in the Sky: Improving Urban-Scene Segmentation via Height-driven Attention Networks
Sungha Choi, Joanne T. Kim, Jaegul Choo
cs.CVarXiv:2003.05128v32020Graph Structure of Neural Networks
Jiaxuan You, Jure Leskovec, Kaiming He +1
cs.LGcs.CVcs.SIarXiv:2007.06559v22020Localizing Object-level Shape Variations with Text-to-Image Diffusion Models
Or Patashnik, Daniel Garibi, Idan Azuri +2
cs.CVcs.GRcs.LGarXiv:2303.11306v22023HeadGAN: One-shot Neural Head Synthesis and Editing
Michail Christos Doukas, Stefanos Zafeiriou, Viktoriia Sharmanska
cs.CVarXiv:2012.08261v32020Contrastive Knowledge Distillation for Anomaly Detection in Multi-Illumination/Focus Display Images
Jihyun Lee, Hangil Park, Yongmin Seo +4
cs.CVcs.LGarXiv:2609.05520v12026CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs
Nhat-Tan Bui, Varshini Elangovan, Arun Reddy Anugu +7
cs.CVcs.LGarXiv:2609.08345v12026AlphaPilot: Autonomous Drone Racing
Philipp Foehn, Dario Brescianini, Elia Kaufmann +4
cs.ROcs.CVeess.SYarXiv:2005.12813v22020Do Depressive Facial Patterns Transfer Across Cultures and Contexts? Evidence from a German RCT and E-DAIC
Misha Sadeghi, Robert Richer, Lydia Helene Rupp +7
cs.CVcs.HCarXiv:2609.05543v12026Exploiting Recurrent Neural Networks and Leap Motion Controller for Sign Language and Semaphoric Gesture Recognition
Danilo Avola, Marco Bernardi, Luigi Cinque +2
cs.CVarXiv:1803.10435v12018Universal Humanoid Motion Representations for Physics-Based Control
Zhengyi Luo, Jinkun Cao, Josh Merel +4
cs.CVcs.GRcs.ROarXiv:2310.04582v22023Situation Awareness for Intelligent Data Distribution in Connected Vehicles
Falk Dettinger, Akshay Narla, Michael Weyrich
cs.ROcs.AIcs.CVarXiv:2609.05521v12026Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through Estimation
Zechun Liu, Kwang-Ting Cheng, Dong Huang +2
cs.CVcs.AIcs.LGarXiv:2111.14826v22021Representation learning of human cortical folding to reveal long lasting neurodevelopmental signatures
Julien Laval, Robin Guiavarch, Antoine Dufournet +24
q-bio.QMcs.CVcs.LGarXiv:2609.05438v12026Coarse-to-Fine Vision-Language Pre-training with Fusion in the Backbone
Zi-Yi Dou, Aishwarya Kamath, Zhe Gan +9
cs.CVcs.CLcs.LGarXiv:2206.07643v22022Efficient Learning on Point Clouds with Basis Point Sets
Sergey Prokudin, Christoph Lassner, Javier Romero
cs.CVarXiv:1908.09186v12019SelfOcc: Self-Supervised Vision-Based 3D Occupancy Prediction
Yuanhui Huang, Wenzhao Zheng, Borui Zhang +2
cs.CVcs.AIcs.LGarXiv:2311.12754v22023FCA: Learning a 3D Full-coverage Vehicle Camouflage for Multi-view Physical Adversarial Attack
Donghua Wang, Tingsong Jiang, Jialiang Sun +5
cs.CVcs.AIarXiv:2109.07193v32021SurgicalSAM: Efficient Class Promptable Surgical Instrument Segmentation
Wenxi Yue, Jing Zhang, Kun Hu +3
cs.CVcs.AIcs.ROarXiv:2308.08746v22023Perceptual Quality Assessment of Omnidirectional Images
Huiyu Duan, Guangtao Zhai, Xiongkuo Min +3
cs.CVarXiv:2207.02674v12022Streamlined Dense Video Captioning
Jonghwan Mun, Linjie Yang, Zhou Ren +2
cs.CVarXiv:1904.03870v12019A Deep Pyramid Deformable Part Model for Face Detection
Rajeev Ranjan, Vishal M. Patel, Rama Chellappa
cs.CVarXiv:1508.04389v12015Video Cloze Procedure for Self-Supervised Spatio-Temporal Learning
Dezhao Luo, Chang Liu, Yu Zhou +4
cs.CVarXiv:2001.00294v12020rPPG-Toolbox: Deep Remote PPG Toolbox
Xin Liu, Girish Narayanswamy, Akshay Paruchuri +7
cs.CVarXiv:2210.00716v32022On Learning 3D Face Morphable Model from In-the-wild Images
Luan Tran, Xiaoming Liu
cs.CVarXiv:1808.09560v22018