Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
14,641 to 14,700 of 18,866
Deep Unfolding Network for Image Super-Resolution
Kai Zhang, Luc Van Gool, Radu Timofte
eess.IVcs.CVarXiv:2003.10428v12020Rotational Projection Statistics for 3D Local Surface Description and Object Recognition
Yulan Guo, Ferdous Sohel, Mohammed Bennamoun +2
cs.CVarXiv:1304.3192v12013Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality
Tristan Thrush, Ryan Jiang, Max Bartolo +4
cs.CVcs.CLarXiv:2204.03162v22022Siamese Instance Search for Tracking
Ran Tao, Efstratios Gavves, Arnold W. M. Smeulders
cs.CVarXiv:1605.05863v12016WorldBench: Benchmarking Physical Understanding of World Models by Isolating Physics Concepts
Rishi Upadhyay, Howard Zhang, Jim Solomon +5
cs.CVarXiv:2601.21282v22026Physics-guided Neural Networks (PGNN): An Application in Lake Temperature Modeling
Arka Daw, Anuj Karpatne, William Watkins +2
cs.LGcs.AIcs.CVarXiv:1710.11431v32017SynthSeg: Segmentation of brain MRI scans of any contrast and resolution without retraining
Benjamin Billot, Douglas N. Greve, Oula Puonti +5
eess.IVcs.CVarXiv:2107.09559v42021Mind the Student: Behavioral and Contextual Cues for Automated Engagement Prediction in Online Learning
Alperen Kantarci, Visvanathan Ramesh, Gemma Roig
cs.CVcs.AIcs.HCarXiv:2608.24340v12026ExpAlign: Expectation-Guided Vision-Language Alignment for Open-Vocabulary Grounding
Junyi Hu, Tian Bai, Fengyi Wu +3
cs.CVarXiv:2601.22666v12026nocaps: novel object captioning at scale
Harsh Agrawal, Karan Desai, Yufei Wang +7
cs.CVcs.AIcs.CLarXiv:1812.08658v32018Zero-Shot Text-Guided Object Generation with Dream Fields
Ajay Jain, Ben Mildenhall, Jonathan T. Barron +2
cs.CVcs.AIcs.GRarXiv:2112.01455v22021Generating High-Quality Crowd Density Maps using Contextual Pyramid CNNs
Vishwanath A. Sindagi, Vishal M. Patel
cs.CVarXiv:1708.00953v12017ObjEmbed: Towards Universal Multimodal Object Embeddings
Shenghao Fu, Yukun Su, Fengyun Rao +3
cs.CVarXiv:2602.01753v32026Drone-based RGB-Infrared Cross-Modality Vehicle Detection via Uncertainty-Aware Learning
Yiming Sun, Bing Cao, Pengfei Zhu +1
cs.CVcs.LGeess.IVarXiv:2003.02437v22020Siam R-CNN: Visual Tracking by Re-Detection
Paul Voigtlaender, Jonathon Luiten, Philip H. S. Torr +1
cs.CVarXiv:1911.12836v22019FOTBCD: A Large-Scale Building Change Detection Benchmark from French Orthophotos and Topographic Data
Abdelrrahman Moubane
cs.CVarXiv:2601.22596v12026Modular Primitives for High-Performance Differentiable Rendering
Samuli Laine, Janne Hellsten, Tero Karras +3
cs.GRcs.CVcs.LGarXiv:2011.03277v12020Semantic Image Segmentation via Deep Parsing Network
Ziwei Liu, Xiaoxiao Li, Ping Luo +2
cs.CVarXiv:1509.02634v22015Generative Adversarial Networks for Extreme Learned Image Compression
Eirikur Agustsson, Michael Tschannen, Fabian Mentzer +2
cs.CVcs.LGarXiv:1804.02958v32018Invertible Conditional GANs for image editing
Guim Perarnau, Joost van de Weijer, Bogdan Raducanu +1
cs.CVcs.AIarXiv:1611.06355v12016NativeTok: Native Visual Tokenization for Improved Image Generation
Bin Wu, Mengqi Huang, Weinan Jia +1
cs.CVarXiv:2601.22837v12026The Deepfake Detection Challenge (DFDC) Preview Dataset
Brian Dolhansky, Russ Howes, Ben Pflaum +2
cs.CVcs.CYarXiv:1910.08854v22019IDeaL: Data-Free Multi-Teacher Distillation via Improved Dead Leaves
Feyza Yavuz, Mert Bülent Sarıyıldız, Diane Larlus
cs.CVarXiv:2608.24759v12026Parabolic Position Encoding: Vision-Centric, Principled, Extrapolatable, General
Christoffer Koo Øhrstrøm, Rafael I. Cabral Muchacho, Yifei Dong +4
cs.CVcs.LGarXiv:2602.01418v22026Luce: Relightable Gaussians for 3D Asset Generation
Mayank Singh, Michele Stoppa, Alvise Memo +7
cs.CVcs.AIcs.GRarXiv:2608.23943v12026BCN20000: Dermoscopic Lesions in the Wild
Marc Combalia, Noel C. F. Codella, Veronica Rotemberg +8
eess.IVcs.CVarXiv:1908.02288v22019Implicit neural representation of textures
Albert Kwok, Zheyuan Hu, Dounia Hammou
cs.CVcs.AIcs.GRarXiv:2602.02354v12026Knockoff Nets: Stealing Functionality of Black-Box Models
Tribhuvanesh Orekondy, Bernt Schiele, Mario Fritz
cs.CVcs.CRcs.LGarXiv:1812.02766v12018EgoActor: Grounding Task Planning into Spatial-aware Egocentric Actions for Humanoid Robots via Visual-Language Models
Yu Bai, MingMing Yu, Chaojie Li +3
cs.ROcs.CVarXiv:2602.04515v12026Segment Anything in High Quality
Lei Ke, Mingqiao Ye, Martin Danelljan +4
cs.CVarXiv:2306.01567v22023LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation
Bo Miao, Weijia Liu, Jun Luo +8
cs.CVcs.ROarXiv:2602.02220v22026DETRs with Collaborative Hybrid Assignments Training
Zhuofan Zong, Guanglu Song, Yu Liu
cs.CVarXiv:2211.12860v62022XCiT: Cross-Covariance Image Transformers
Alaaeldin El-Nouby, Hugo Touvron, Mathilde Caron +8
cs.CVcs.LGarXiv:2106.09681v22021Masked Autoencoders As Spatiotemporal Learners
Christoph Feichtenhofer, Haoqi Fan, Yanghao Li +1
cs.CVcs.LGarXiv:2205.09113v22022DeiT III: Revenge of the ViT
Hugo Touvron, Matthieu Cord, Hervé Jégou
cs.CVarXiv:2204.07118v12022Source-Face Authenticity Detection for 3D Gaussian Heads Reconstructed from a Single Portrait: A Benchmark and Dedicated Detector
Yujie Gao, Zijian Yu, Yan Hong +2
cs.CVarXiv:2608.23984v12026RANKVIDEO: Reasoning Reranking for Text-to-Video Retrieval
Tyler Skow, Alexander Martin, Benjamin Van Durme +2
cs.IRcs.CVarXiv:2602.02444v22026Group-based Sparse Representation for Image Restoration
Jian Zhang, Debin Zhao, Wen Gao
cs.CVarXiv:1405.3351v12014FaceLinkGen: Rethinking Identity Leakage in Privacy-Preserving Face Recognition with Identity Extraction
Wenqi Guo, Shan Du
cs.CVarXiv:2602.02914v22026EgoErrorVQA: Assess Egocentric Comprehension Capabilities through Procedural Errors for Ego-Agentic AI
Junlong Li, Junxi Li, Jianjun Gao +3
cs.CVarXiv:2608.24134v12026SkeletonGaussian: Editable 4D Generation through Gaussian Skeletonization
Lifan Wu, Ruijie Zhu, Yubo Ai +1
cs.CVcs.AIcs.GRarXiv:2602.04271v12026Real-World Knowledge-Guided Change Data Synthesis for Remote Sensing
Yaoyi Qi, Xingxing Weng, Chao Pang +5
cs.AIcs.CVarXiv:2608.24263v12026DM-GAN: Dynamic Memory Generative Adversarial Networks for Text-to-Image Synthesis
Minfeng Zhu, Pingbo Pan, Wei Chen +1
cs.CVarXiv:1904.01310v12019Distilling Knowledge via Knowledge Review
Pengguang Chen, Shu Liu, Hengshuang Zhao +1
cs.CVarXiv:2104.09044v12021BasicVSR: The Search for Essential Components in Video Super-Resolution and Beyond
Kelvin C. K. Chan, Xintao Wang, Ke Yu +2
cs.CVarXiv:2012.02181v22020Detecting Cancer Metastases on Gigapixel Pathology Images
Yun Liu, Krishna Gadepalli, Mohammad Norouzi +10
cs.CVarXiv:1703.02442v22017Axial Attention in Multidimensional Transformers
Jonathan Ho, Nal Kalchbrenner, Dirk Weissenborn +1
cs.CVarXiv:1912.12180v12019End-to-End Semi-Supervised Object Detection with Soft Teacher
Mengde Xu, Zheng Zhang, Han Hu +5
cs.CVcs.AIarXiv:2106.09018v32021RepViT: Revisiting Mobile CNN From ViT Perspective
Ao Wang, Hui Chen, Zijia Lin +2
cs.CVarXiv:2307.09283v82023Deep-Emotion: Facial Expression Recognition Using Attentional Convolutional Network
Shervin Minaee, Amirali Abdolrashidi
cs.CVarXiv:1902.01019v12019Beyond Face Rotation: Global and Local Perception GAN for Photorealistic and Identity Preserving Frontal View Synthesis
Rui Huang, Shu Zhang, Tianyu Li +1
cs.CVarXiv:1704.04086v22017Multiresolution Knowledge Distillation for Anomaly Detection
Mohammadreza Salehi, Niousha Sadjadi, Soroosh Baselizadeh +2
cs.CVarXiv:2011.11108v12020ZODIAC: Zero-shot Octree-based Diffusion for Anatomical Completion
Miruna-Alexandra Gafencu, Vlad Bratulescu, Yordanka Velikova +2
cs.CVarXiv:2608.24422v12026The ApolloScape Open Dataset for Autonomous Driving and its Application
Xinyu Huang, Peng Wang, Xinjing Cheng +3
cs.CVarXiv:1803.06184v42018Your Classifier is Secretly an Energy Based Model and You Should Treat it Like One
Will Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen +3
cs.LGcs.CVstat.MLarXiv:1912.03263v32019Rethinking Global Text Conditioning in Diffusion Transformers
Nikita Starodubcev, Daniil Pakhomov, Zongze Wu +6
cs.CVarXiv:2602.09268v12026What Does Prompt Learning Change? -A Natural-Language Concept Analysis of Vision-Language Models
Ryo Kamiya, Hiroshi Kera, Kazuhiko Kawamoto
cs.CVarXiv:2608.24142v12026Automatic Liver and Lesion Segmentation in CT Using Cascaded Fully Convolutional Neural Networks and 3D Conditional Random Fields
Patrick Ferdinand Christ, Mohamed Ezzeldin A. Elshaer, Florian Ettlinger +10
cs.CVarXiv:1610.02177v12016Generate To Adapt: Aligning Domains using Generative Adversarial Networks
Swami Sankaranarayanan, Yogesh Balaji, Carlos D. Castillo +1
cs.CVarXiv:1704.01705v42017Best of Both Worlds: Multimodal Reasoning and Generation via Unified Discrete Flow Matching
Onkar Susladkar, Tushar Prakash, Gayatri Deshmukh +8
cs.CVarXiv:2602.12221v22026