Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
14,581 to 14,640 of 18,830
ROI-Gated SAHI: Content-Adaptive Slicing-Based Inference for Efficient Object Detection
Rashid Riyadh, Abd Ullah Khan, Imad Gohar +1
cs.CVarXiv:2608.23923v12026Slimmable Neural Networks
Jiahui Yu, Linjie Yang, Ning Xu +2
cs.CVcs.AIarXiv:1812.08928v12018SketchJudge: A Diagnostic Benchmark for Grading Hand-drawn Diagrams with Multimodal Large Language Models
Yuhang Su, Mei Wang, Yaoyao Zhong +4
cs.CVcs.AIarXiv:2601.06944v12026An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Liang Chen, Haozhe Zhao, Tianyu Liu +4
cs.CVcs.AIcs.CLarXiv:2403.06764v32024GeoMotionGPT: Geometry-Aligned Motion Understanding with Large Language Models
Zhankai Ye, Bofan Li, Yukai Jin +5
cs.CVcs.AIarXiv:2601.07632v42026CurricularFace: Adaptive Curriculum Learning Loss for Deep Face Recognition
Yuge Huang, Yuhan Wang, Ying Tai +5
cs.CVarXiv:2004.00288v12020KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning
Egor Cherepanov, Daniil Zelezetsky, Alexey K. Kovalev +1
cs.LGcs.AIcs.CVarXiv:2601.14232v22026Large Multimodal Models as General In-Context Classifiers
Marco Garosi, Matteo Farina, Alessandro Conti +2
cs.CVarXiv:2602.23229v12026Alterbute: Editing Intrinsic Attributes of Objects in Images
Tal Reiss, Daniel Winter, Matan Cohen +4
cs.CVcs.GRarXiv:2601.10714v22026CLEVRER: CoLlision Events for Video REpresentation and Reasoning
Kexin Yi, Chuang Gan, Yunzhu Li +4
cs.CVcs.AIcs.CLarXiv:1910.01442v22019Mutual Mean-Teaching: Pseudo Label Refinery for Unsupervised Domain Adaptation on Person Re-identification
Yixiao Ge, Dapeng Chen, Hongsheng Li
cs.CVarXiv:2001.01526v22020Multi30K: Multilingual English-German Image Descriptions
Desmond Elliott, Stella Frank, Khalil Sima'an +1
cs.CLcs.CVarXiv:1605.00459v12016PredRNN: A Recurrent Neural Network for Spatiotemporal Predictive Learning
Yunbo Wang, Haixu Wu, Jianjin Zhang +4
cs.LGcs.CVarXiv:2103.09504v42021Dynamic Neural Radiance Fields for Monocular 4D Facial Avatar Reconstruction
Guy Gafni, Justus Thies, Michael Zollhöfer +1
cs.CVcs.GRarXiv:2012.03065v12020Low-Rank Ternary Adaptation for Fine-Tuning Transformers
Alexandru-Dragos Manolache, Yunqiang Li, Jan van Gemert
cs.CVcs.LGarXiv:2608.24469v12026CenterMask : Real-Time Anchor-Free Instance Segmentation
Youngwan Lee, Jongyoul Park
cs.CVarXiv:1911.06667v62019MOTS: Multi-Object Tracking and Segmentation
Paul Voigtlaender, Michael Krause, Aljosa Osep +4
cs.CVarXiv:1902.03604v22019CA-Net: Comprehensive Attention Convolutional Neural Networks for Explainable Medical Image Segmentation
Ran Gu, Guotai Wang, Tao Song +6
eess.IVcs.CVarXiv:2009.10549v22020Learning to Upsample by Learning to Sample
Wenze Liu, Hao Lu, Hongtao Fu +1
cs.CVarXiv:2308.15085v12023RemoteVAR: Autoregressive Visual Modeling for Remote Sensing Change Detection
Yilmaz Korkmaz, Vishal M. Patel
cs.CVarXiv:2601.11898v12026Word-level Deep Sign Language Recognition from Video: A New Large-scale Dataset and Methods Comparison
Dongxu Li, Cristian Rodriguez Opazo, Xin Yu +1
cs.CVcs.HCcs.MMarXiv:1910.11006v22019DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
Zhiyu Wu, Xiaokang Chen, Zizheng Pan +24
cs.CVcs.AIcs.CLarXiv:2412.10302v12024Learning a Discriminative Null Space for Person Re-identification
Li Zhang, Tao Xiang, Shaogang Gong
cs.CVarXiv:1603.02139v12016AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed Gradients
Juntang Zhuang, Tommy Tang, Yifan Ding +4
cs.LGcs.CVstat.MLarXiv:2010.07468v52020Deep Unfolding Network for Image Super-Resolution
Kai Zhang, Luc Van Gool, Radu Timofte
eess.IVcs.CVarXiv:2003.10428v12020Rotational Projection Statistics for 3D Local Surface Description and Object Recognition
Yulan Guo, Ferdous Sohel, Mohammed Bennamoun +2
cs.CVarXiv:1304.3192v12013Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality
Tristan Thrush, Ryan Jiang, Max Bartolo +4
cs.CVcs.CLarXiv:2204.03162v22022Siamese Instance Search for Tracking
Ran Tao, Efstratios Gavves, Arnold W. M. Smeulders
cs.CVarXiv:1605.05863v12016WorldBench: Benchmarking Physical Understanding of World Models by Isolating Physics Concepts
Rishi Upadhyay, Howard Zhang, Jim Solomon +5
cs.CVarXiv:2601.21282v22026Physics-guided Neural Networks (PGNN): An Application in Lake Temperature Modeling
Arka Daw, Anuj Karpatne, William Watkins +2
cs.LGcs.AIcs.CVarXiv:1710.11431v32017SynthSeg: Segmentation of brain MRI scans of any contrast and resolution without retraining
Benjamin Billot, Douglas N. Greve, Oula Puonti +5
eess.IVcs.CVarXiv:2107.09559v42021Mind the Student: Behavioral and Contextual Cues for Automated Engagement Prediction in Online Learning
Alperen Kantarci, Visvanathan Ramesh, Gemma Roig
cs.CVcs.AIcs.HCarXiv:2608.24340v12026ExpAlign: Expectation-Guided Vision-Language Alignment for Open-Vocabulary Grounding
Junyi Hu, Tian Bai, Fengyi Wu +3
cs.CVarXiv:2601.22666v12026nocaps: novel object captioning at scale
Harsh Agrawal, Karan Desai, Yufei Wang +7
cs.CVcs.AIcs.CLarXiv:1812.08658v32018Zero-Shot Text-Guided Object Generation with Dream Fields
Ajay Jain, Ben Mildenhall, Jonathan T. Barron +2
cs.CVcs.AIcs.GRarXiv:2112.01455v22021Generating High-Quality Crowd Density Maps using Contextual Pyramid CNNs
Vishwanath A. Sindagi, Vishal M. Patel
cs.CVarXiv:1708.00953v12017ObjEmbed: Towards Universal Multimodal Object Embeddings
Shenghao Fu, Yukun Su, Fengyun Rao +3
cs.CVarXiv:2602.01753v32026Drone-based RGB-Infrared Cross-Modality Vehicle Detection via Uncertainty-Aware Learning
Yiming Sun, Bing Cao, Pengfei Zhu +1
cs.CVcs.LGeess.IVarXiv:2003.02437v22020Siam R-CNN: Visual Tracking by Re-Detection
Paul Voigtlaender, Jonathon Luiten, Philip H. S. Torr +1
cs.CVarXiv:1911.12836v22019FOTBCD: A Large-Scale Building Change Detection Benchmark from French Orthophotos and Topographic Data
Abdelrrahman Moubane
cs.CVarXiv:2601.22596v12026Modular Primitives for High-Performance Differentiable Rendering
Samuli Laine, Janne Hellsten, Tero Karras +3
cs.GRcs.CVcs.LGarXiv:2011.03277v12020Semantic Image Segmentation via Deep Parsing Network
Ziwei Liu, Xiaoxiao Li, Ping Luo +2
cs.CVarXiv:1509.02634v22015Generative Adversarial Networks for Extreme Learned Image Compression
Eirikur Agustsson, Michael Tschannen, Fabian Mentzer +2
cs.CVcs.LGarXiv:1804.02958v32018Invertible Conditional GANs for image editing
Guim Perarnau, Joost van de Weijer, Bogdan Raducanu +1
cs.CVcs.AIarXiv:1611.06355v12016NativeTok: Native Visual Tokenization for Improved Image Generation
Bin Wu, Mengqi Huang, Weinan Jia +1
cs.CVarXiv:2601.22837v12026The Deepfake Detection Challenge (DFDC) Preview Dataset
Brian Dolhansky, Russ Howes, Ben Pflaum +2
cs.CVcs.CYarXiv:1910.08854v22019IDeaL: Data-Free Multi-Teacher Distillation via Improved Dead Leaves
Feyza Yavuz, Mert Bülent Sarıyıldız, Diane Larlus
cs.CVarXiv:2608.24759v12026Parabolic Position Encoding: Vision-Centric, Principled, Extrapolatable, General
Christoffer Koo Øhrstrøm, Rafael I. Cabral Muchacho, Yifei Dong +4
cs.CVcs.LGarXiv:2602.01418v22026Luce: Relightable Gaussians for 3D Asset Generation
Mayank Singh, Michele Stoppa, Alvise Memo +7
cs.CVcs.AIcs.GRarXiv:2608.23943v12026BCN20000: Dermoscopic Lesions in the Wild
Marc Combalia, Noel C. F. Codella, Veronica Rotemberg +8
eess.IVcs.CVarXiv:1908.02288v22019Implicit neural representation of textures
Albert Kwok, Zheyuan Hu, Dounia Hammou
cs.CVcs.AIcs.GRarXiv:2602.02354v12026Knockoff Nets: Stealing Functionality of Black-Box Models
Tribhuvanesh Orekondy, Bernt Schiele, Mario Fritz
cs.CVcs.CRcs.LGarXiv:1812.02766v12018EgoActor: Grounding Task Planning into Spatial-aware Egocentric Actions for Humanoid Robots via Visual-Language Models
Yu Bai, MingMing Yu, Chaojie Li +3
cs.ROcs.CVarXiv:2602.04515v12026Segment Anything in High Quality
Lei Ke, Mingqiao Ye, Martin Danelljan +4
cs.CVarXiv:2306.01567v22023LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation
Bo Miao, Weijia Liu, Jun Luo +8
cs.CVcs.ROarXiv:2602.02220v22026DETRs with Collaborative Hybrid Assignments Training
Zhuofan Zong, Guanglu Song, Yu Liu
cs.CVarXiv:2211.12860v62022XCiT: Cross-Covariance Image Transformers
Alaaeldin El-Nouby, Hugo Touvron, Mathilde Caron +8
cs.CVcs.LGarXiv:2106.09681v22021Masked Autoencoders As Spatiotemporal Learners
Christoph Feichtenhofer, Haoqi Fan, Yanghao Li +1
cs.CVcs.LGarXiv:2205.09113v22022DeiT III: Revenge of the ViT
Hugo Touvron, Matthieu Cord, Hervé Jégou
cs.CVarXiv:2204.07118v12022Source-Face Authenticity Detection for 3D Gaussian Heads Reconstructed from a Single Portrait: A Benchmark and Dedicated Detector
Yujie Gao, Zijian Yu, Yan Hong +2
cs.CVarXiv:2608.23984v12026