Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

14,641 to 14,700 of 18,866

  1. Deep Unfolding Network for Image Super-Resolution

    Kai Zhang, Luc Van Gool, Radu Timofte

    eess.IVcs.CVarXiv:2003.10428v12020
  2. Rotational Projection Statistics for 3D Local Surface Description and Object Recognition

    Yulan Guo, Ferdous Sohel, Mohammed Bennamoun +2

    cs.CVarXiv:1304.3192v12013
  3. Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality

    Tristan Thrush, Ryan Jiang, Max Bartolo +4

    cs.CVcs.CLarXiv:2204.03162v22022
  4. Siamese Instance Search for Tracking

    Ran Tao, Efstratios Gavves, Arnold W. M. Smeulders

    cs.CVarXiv:1605.05863v12016
  5. WorldBench: Benchmarking Physical Understanding of World Models by Isolating Physics Concepts

    Rishi Upadhyay, Howard Zhang, Jim Solomon +5

    cs.CVarXiv:2601.21282v22026
  6. Physics-guided Neural Networks (PGNN): An Application in Lake Temperature Modeling

    Arka Daw, Anuj Karpatne, William Watkins +2

    cs.LGcs.AIcs.CVarXiv:1710.11431v32017
  7. SynthSeg: Segmentation of brain MRI scans of any contrast and resolution without retraining

    Benjamin Billot, Douglas N. Greve, Oula Puonti +5

    eess.IVcs.CVarXiv:2107.09559v42021
  8. Mind the Student: Behavioral and Contextual Cues for Automated Engagement Prediction in Online Learning

    Alperen Kantarci, Visvanathan Ramesh, Gemma Roig

    cs.CVcs.AIcs.HCarXiv:2608.24340v12026
  9. ExpAlign: Expectation-Guided Vision-Language Alignment for Open-Vocabulary Grounding

    Junyi Hu, Tian Bai, Fengyi Wu +3

    cs.CVarXiv:2601.22666v12026
  10. nocaps: novel object captioning at scale

    Harsh Agrawal, Karan Desai, Yufei Wang +7

    cs.CVcs.AIcs.CLarXiv:1812.08658v32018
  11. Zero-Shot Text-Guided Object Generation with Dream Fields

    Ajay Jain, Ben Mildenhall, Jonathan T. Barron +2

    cs.CVcs.AIcs.GRarXiv:2112.01455v22021
  12. Generating High-Quality Crowd Density Maps using Contextual Pyramid CNNs

    Vishwanath A. Sindagi, Vishal M. Patel

    cs.CVarXiv:1708.00953v12017
  13. ObjEmbed: Towards Universal Multimodal Object Embeddings

    Shenghao Fu, Yukun Su, Fengyun Rao +3

    cs.CVarXiv:2602.01753v32026
  14. Drone-based RGB-Infrared Cross-Modality Vehicle Detection via Uncertainty-Aware Learning

    Yiming Sun, Bing Cao, Pengfei Zhu +1

    cs.CVcs.LGeess.IVarXiv:2003.02437v22020
  15. Siam R-CNN: Visual Tracking by Re-Detection

    Paul Voigtlaender, Jonathon Luiten, Philip H. S. Torr +1

    cs.CVarXiv:1911.12836v22019
  16. FOTBCD: A Large-Scale Building Change Detection Benchmark from French Orthophotos and Topographic Data

    Abdelrrahman Moubane

    cs.CVarXiv:2601.22596v12026
  17. Modular Primitives for High-Performance Differentiable Rendering

    Samuli Laine, Janne Hellsten, Tero Karras +3

    cs.GRcs.CVcs.LGarXiv:2011.03277v12020
  18. Semantic Image Segmentation via Deep Parsing Network

    Ziwei Liu, Xiaoxiao Li, Ping Luo +2

    cs.CVarXiv:1509.02634v22015
  19. Generative Adversarial Networks for Extreme Learned Image Compression

    Eirikur Agustsson, Michael Tschannen, Fabian Mentzer +2

    cs.CVcs.LGarXiv:1804.02958v32018
  20. Invertible Conditional GANs for image editing

    Guim Perarnau, Joost van de Weijer, Bogdan Raducanu +1

    cs.CVcs.AIarXiv:1611.06355v12016
  21. NativeTok: Native Visual Tokenization for Improved Image Generation

    Bin Wu, Mengqi Huang, Weinan Jia +1

    cs.CVarXiv:2601.22837v12026
  22. The Deepfake Detection Challenge (DFDC) Preview Dataset

    Brian Dolhansky, Russ Howes, Ben Pflaum +2

    cs.CVcs.CYarXiv:1910.08854v22019
  23. IDeaL: Data-Free Multi-Teacher Distillation via Improved Dead Leaves

    Feyza Yavuz, Mert Bülent Sarıyıldız, Diane Larlus

    cs.CVarXiv:2608.24759v12026
  24. Parabolic Position Encoding: Vision-Centric, Principled, Extrapolatable, General

    Christoffer Koo Øhrstrøm, Rafael I. Cabral Muchacho, Yifei Dong +4

    cs.CVcs.LGarXiv:2602.01418v22026
  25. Luce: Relightable Gaussians for 3D Asset Generation

    Mayank Singh, Michele Stoppa, Alvise Memo +7

    cs.CVcs.AIcs.GRarXiv:2608.23943v12026
  26. BCN20000: Dermoscopic Lesions in the Wild

    Marc Combalia, Noel C. F. Codella, Veronica Rotemberg +8

    eess.IVcs.CVarXiv:1908.02288v22019
  27. Implicit neural representation of textures

    Albert Kwok, Zheyuan Hu, Dounia Hammou

    cs.CVcs.AIcs.GRarXiv:2602.02354v12026
  28. Knockoff Nets: Stealing Functionality of Black-Box Models

    Tribhuvanesh Orekondy, Bernt Schiele, Mario Fritz

    cs.CVcs.CRcs.LGarXiv:1812.02766v12018
  29. EgoActor: Grounding Task Planning into Spatial-aware Egocentric Actions for Humanoid Robots via Visual-Language Models

    Yu Bai, MingMing Yu, Chaojie Li +3

    cs.ROcs.CVarXiv:2602.04515v12026
  30. Segment Anything in High Quality

    Lei Ke, Mingqiao Ye, Martin Danelljan +4

    cs.CVarXiv:2306.01567v22023
  31. LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation

    Bo Miao, Weijia Liu, Jun Luo +8

    cs.CVcs.ROarXiv:2602.02220v22026
  32. DETRs with Collaborative Hybrid Assignments Training

    Zhuofan Zong, Guanglu Song, Yu Liu

    cs.CVarXiv:2211.12860v62022
  33. XCiT: Cross-Covariance Image Transformers

    Alaaeldin El-Nouby, Hugo Touvron, Mathilde Caron +8

    cs.CVcs.LGarXiv:2106.09681v22021
  34. Masked Autoencoders As Spatiotemporal Learners

    Christoph Feichtenhofer, Haoqi Fan, Yanghao Li +1

    cs.CVcs.LGarXiv:2205.09113v22022
  35. DeiT III: Revenge of the ViT

    Hugo Touvron, Matthieu Cord, Hervé Jégou

    cs.CVarXiv:2204.07118v12022
  36. Source-Face Authenticity Detection for 3D Gaussian Heads Reconstructed from a Single Portrait: A Benchmark and Dedicated Detector

    Yujie Gao, Zijian Yu, Yan Hong +2

    cs.CVarXiv:2608.23984v12026
  37. RANKVIDEO: Reasoning Reranking for Text-to-Video Retrieval

    Tyler Skow, Alexander Martin, Benjamin Van Durme +2

    cs.IRcs.CVarXiv:2602.02444v22026
  38. Group-based Sparse Representation for Image Restoration

    Jian Zhang, Debin Zhao, Wen Gao

    cs.CVarXiv:1405.3351v12014
  39. FaceLinkGen: Rethinking Identity Leakage in Privacy-Preserving Face Recognition with Identity Extraction

    Wenqi Guo, Shan Du

    cs.CVarXiv:2602.02914v22026
  40. EgoErrorVQA: Assess Egocentric Comprehension Capabilities through Procedural Errors for Ego-Agentic AI

    Junlong Li, Junxi Li, Jianjun Gao +3

    cs.CVarXiv:2608.24134v12026
  41. SkeletonGaussian: Editable 4D Generation through Gaussian Skeletonization

    Lifan Wu, Ruijie Zhu, Yubo Ai +1

    cs.CVcs.AIcs.GRarXiv:2602.04271v12026
  42. Real-World Knowledge-Guided Change Data Synthesis for Remote Sensing

    Yaoyi Qi, Xingxing Weng, Chao Pang +5

    cs.AIcs.CVarXiv:2608.24263v12026
  43. DM-GAN: Dynamic Memory Generative Adversarial Networks for Text-to-Image Synthesis

    Minfeng Zhu, Pingbo Pan, Wei Chen +1

    cs.CVarXiv:1904.01310v12019
  44. Distilling Knowledge via Knowledge Review

    Pengguang Chen, Shu Liu, Hengshuang Zhao +1

    cs.CVarXiv:2104.09044v12021
  45. BasicVSR: The Search for Essential Components in Video Super-Resolution and Beyond

    Kelvin C. K. Chan, Xintao Wang, Ke Yu +2

    cs.CVarXiv:2012.02181v22020
  46. Detecting Cancer Metastases on Gigapixel Pathology Images

    Yun Liu, Krishna Gadepalli, Mohammad Norouzi +10

    cs.CVarXiv:1703.02442v22017
  47. Axial Attention in Multidimensional Transformers

    Jonathan Ho, Nal Kalchbrenner, Dirk Weissenborn +1

    cs.CVarXiv:1912.12180v12019
  48. End-to-End Semi-Supervised Object Detection with Soft Teacher

    Mengde Xu, Zheng Zhang, Han Hu +5

    cs.CVcs.AIarXiv:2106.09018v32021
  49. RepViT: Revisiting Mobile CNN From ViT Perspective

    Ao Wang, Hui Chen, Zijia Lin +2

    cs.CVarXiv:2307.09283v82023
  50. Deep-Emotion: Facial Expression Recognition Using Attentional Convolutional Network

    Shervin Minaee, Amirali Abdolrashidi

    cs.CVarXiv:1902.01019v12019
  51. Beyond Face Rotation: Global and Local Perception GAN for Photorealistic and Identity Preserving Frontal View Synthesis

    Rui Huang, Shu Zhang, Tianyu Li +1

    cs.CVarXiv:1704.04086v22017
  52. Multiresolution Knowledge Distillation for Anomaly Detection

    Mohammadreza Salehi, Niousha Sadjadi, Soroosh Baselizadeh +2

    cs.CVarXiv:2011.11108v12020
  53. ZODIAC: Zero-shot Octree-based Diffusion for Anatomical Completion

    Miruna-Alexandra Gafencu, Vlad Bratulescu, Yordanka Velikova +2

    cs.CVarXiv:2608.24422v12026
  54. The ApolloScape Open Dataset for Autonomous Driving and its Application

    Xinyu Huang, Peng Wang, Xinjing Cheng +3

    cs.CVarXiv:1803.06184v42018
  55. Your Classifier is Secretly an Energy Based Model and You Should Treat it Like One

    Will Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen +3

    cs.LGcs.CVstat.MLarXiv:1912.03263v32019
  56. Rethinking Global Text Conditioning in Diffusion Transformers

    Nikita Starodubcev, Daniil Pakhomov, Zongze Wu +6

    cs.CVarXiv:2602.09268v12026
  57. What Does Prompt Learning Change? -A Natural-Language Concept Analysis of Vision-Language Models

    Ryo Kamiya, Hiroshi Kera, Kazuhiko Kawamoto

    cs.CVarXiv:2608.24142v12026
  58. Automatic Liver and Lesion Segmentation in CT Using Cascaded Fully Convolutional Neural Networks and 3D Conditional Random Fields

    Patrick Ferdinand Christ, Mohamed Ezzeldin A. Elshaer, Florian Ettlinger +10

    cs.CVarXiv:1610.02177v12016
  59. Generate To Adapt: Aligning Domains using Generative Adversarial Networks

    Swami Sankaranarayanan, Yogesh Balaji, Carlos D. Castillo +1

    cs.CVarXiv:1704.01705v42017
  60. Best of Both Worlds: Multimodal Reasoning and Generation via Unified Discrete Flow Matching

    Onkar Susladkar, Tushar Prakash, Gayatri Deshmukh +8

    cs.CVarXiv:2602.12221v22026