Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
7,321 to 7,380 of 18,927
Cross-domain Face Presentation Attack Detection via Multi-domain Disentangled Representation Learning
Guoqing Wang, Hu Han, Shiguang Shan +1
cs.CVarXiv:2004.01959v12020Unsupervised Multi-Task Feature Learning on Point Clouds
Kaveh Hassani, Mike Haley
cs.CVcs.LGarXiv:1910.08207v12019PTQD: Accurate Post-Training Quantization for Diffusion Models
Yefei He, Luping Liu, Jing Liu +3
cs.CVarXiv:2305.10657v42023DRAMA: Joint Risk Localization and Captioning in Driving
Srikanth Malla, Chiho Choi, Isht Dwivedi +2
cs.CVcs.AIcs.LGarXiv:2209.10767v22022Conditional Deep Learning for Energy-Efficient and Enhanced Pattern Recognition
Priyadarshini Panda, Abhronil Sengupta, Kaushik Roy
cs.CVarXiv:1509.08971v62015Data-dependent Initializations of Convolutional Neural Networks
Philipp Krähenbühl, Carl Doersch, Jeff Donahue +1
cs.CVcs.LGarXiv:1511.06856v32015Learning a Rotation Invariant Detector with Rotatable Bounding Box
Lei Liu, Zongxu Pan, Bin Lei
cs.CVarXiv:1711.09405v12017Efficient Neighbourhood Consensus Networks via Submanifold Sparse Convolutions
Ignacio Rocco, Relja Arandjelović, Josef Sivic
cs.CVarXiv:2004.10566v12020Epipolar Transformers
Yihui He, Rui Yan, Katerina Fragkiadaki +1
cs.CVarXiv:2005.04551v12020FCN-Transformer Feature Fusion for Polyp Segmentation
Edward Sanderson, Bogdan J. Matuszewski
eess.IVcs.CVcs.LGarXiv:2208.08352v12022An Entropy-based Pruning Method for CNN Compression
Jian-Hao Luo, Jianxin Wu
cs.CVarXiv:1706.05791v12017Panoptic Lifting for 3D Scene Understanding with Neural Fields
Yawar Siddiqui, Lorenzo Porzi, Samuel Rota Buló +4
cs.CVcs.LGarXiv:2212.09802v12022MetaSAug: Meta Semantic Augmentation for Long-Tailed Visual Recognition
Shuang Li, Kaixiong Gong, Chi Harold Liu +3
cs.CVarXiv:2103.12579v32021Principia: Relational Physics Tests for Video Models
Varun Varma Thozhiyoor, Shivam Tripathi, Venkatesh Babu Radhakrishnan +1
cs.CVarXiv:2609.04200v12026FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow
Byeongjun Park, Byung-Hoon Kim, Hyungjin Chung
cs.CVarXiv:2609.03563v12026Supervision-by-Registration: An Unsupervised Approach to Improve the Precision of Facial Landmark Detectors
Xuanyi Dong, Shoou-I Yu, Xinshuo Weng +3
cs.CVarXiv:1807.00966v22018What Would You Expect? Anticipating Egocentric Actions with Rolling-Unrolling LSTMs and Modality Attention
Antonino Furnari, Giovanni Maria Farinella
cs.CVcs.AIarXiv:1905.09035v22019Hierarchical Long Short-Term Concurrent Memory for Human Interaction Recognition
Xiangbo Shu, Jinhui Tang, Guo-Jun Qi +2
cs.CVarXiv:1811.00270v12018Compressing AI Traffic: Standardized Neural Network Coding of Visual-Token Representations in Split Vision-Language Inference
Reza Heidari, Hamed R. Tavakoli, Juho Kannala
cs.CVeess.IVarXiv:2609.01200v12026Motion Guided Attention for Video Salient Object Detection
Haofeng Li, Guanqi Chen, Guanbin Li +1
cs.CVarXiv:1909.07061v22019BS: Take the Hint - Interactive Multitracer PET/CT Lesion Segmentation with a Scribble-Conditioned ResEnc U-Net
Marven Sherif, Amgad Elmasry, Youssef Ghazal +1
cs.CVcs.AIarXiv:2609.01554v12026Grounded Video Description
Luowei Zhou, Yannis Kalantidis, Xinlei Chen +2
cs.CVarXiv:1812.06587v22018PointGPT: Auto-regressively Generative Pre-training from Point Clouds
Guangyan Chen, Meiling Wang, Yi Yang +3
cs.CVarXiv:2305.11487v22023WorldReward: Reward Modeling for Camera-Conditioned World Models
Yibin Wang, Zehan Wang, Junshu Tang +13
cs.CVarXiv:2609.03952v12026The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation
Yichen Liu, Quanwei Zhang, Haozhe Wang +7
cs.MMcs.CVarXiv:2609.02367v12026CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation
Tingyu Song, Mingxin Li, Yanzhao Zhang +5
cs.CVcs.AIcs.CLarXiv:2609.04083v12026GaussianEditor: Editing 3D Gaussians Delicately with Text Instructions
Junjie Wang, Jiemin Fang, Xiaopeng Zhang +2
cs.CVcs.GRarXiv:2311.16037v22023Self-Supervised Learning for Cardiac MR Image Segmentation by Anatomical Position Prediction
Wenjia Bai, Chen Chen, Giacomo Tarroni +6
cs.CVarXiv:1907.02757v12019Sparse Representation-based Open Set Recognition
He Zhang, Vishal M. Patel
cs.CVarXiv:1705.02431v12017Fast Multi-class Dictionaries Learning with Geometrical Directions in MRI Reconstruction
Zhifang Zhan, Jian-Feng Cai, Di Guo +3
cs.CVmath.OCphysics.med-pharXiv:1503.02945v22015ORB-SVM : An Innovative Hybrid Framework for Efficient Brain Tumor Detection from MRI Scans
Amirhosein Azarpour
cs.CVcs.AIarXiv:2609.02333v12026LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes
Chuyan Chen, Haoxing Chen, Kun Chen +27
cs.CVcs.AIarXiv:2609.03796v12026Improved Stereo Matching with Constant Highway Networks and Reflective Confidence Learning
Amit Shaked, Lior Wolf
cs.CVarXiv:1701.00165v12016Just How Toxic is Data Poisoning? A Unified Benchmark for Backdoor and Data Poisoning Attacks
Avi Schwarzschild, Micah Goldblum, Arjun Gupta +2
cs.LGcs.CRcs.CVarXiv:2006.12557v32020TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Liao Qu, Huichao Zhang, Yiheng Liu +7
cs.CVcs.AIarXiv:2412.03069v22024World-Coherent Decoding: Self-Verifying Test-Time Planning for World Action Models
Chuhan Zhang, Seiji Ito, Kenta Hoshino +2
cs.CVarXiv:2609.02159v12026SemanticAdv: Generating Adversarial Examples via Attribute-conditional Image Editing
Haonan Qiu, Chaowei Xiao, Lei Yang +3
cs.LGcs.CRcs.CVarXiv:1906.07927v42019HUGS: Human Gaussian Splats
Muhammed Kocabas, Jen-Hao Rick Chang, James Gabriel +2
cs.CVcs.GRarXiv:2311.17910v12023Editable Free-viewpoint Video Using a Layered Neural Representation
Jiakai Zhang, Xinhang Liu, Xinyi Ye +6
cs.CVcs.GRarXiv:2104.14786v12021Handwriting Trajectory Recovery via Autoregressive Ordered Stroke Instance Prediction
En-Guang Wang, Yan-Ming Zhang, Fei Yin +1
cs.CVarXiv:2609.02251v12026Debugging Tests for Model Explanations
Julius Adebayo, Michael Muelly, Ilaria Liccardi +1
cs.CVcs.LGarXiv:2011.05429v12020The Riemannian Geometry of Deep Generative Models
Hang Shao, Abhishek Kumar, P. Thomas Fletcher
cs.LGcs.CVstat.MLarXiv:1711.08014v12017Contrastive Learning for Image Captioning
Bo Dai, Dahua Lin
cs.CVarXiv:1710.02534v12017LoFi RADIO: A Distilled In-Domain Backbone Applied for Artifact-Severity Grading of Ultra-Low-Field Neonatal Brain MR
Jonathan B. Martin, Yashwant Kurmi, Charlotte R. Sappo
eess.IVcs.CVarXiv:2609.02676v12026Divide-and-Assemble: Learning Block-wise Memory for Unsupervised Anomaly Detection
Jinlei Hou, Yingying Zhang, Qiaoyong Zhong +3
cs.CVarXiv:2107.13118v12021Federated LoRA Adaptation of BiomedCLIP Across Four International Chest X-Ray Cohorts
Sanjaya Poudel, Nirajan Kunwor, Manish Dhakal +2
cs.LGcs.AIcs.CVarXiv:2609.02101v12026Few-Shot Learning via Saliency-guided Hallucination of Samples
Hongguang Zhang, Jing Zhang, Piotr Koniusz
cs.CVarXiv:1904.03472v12019Distance Metric Learning using Graph Convolutional Networks: Application to Functional Brain Networks
Sofia Ira Ktena, Sarah Parisot, Enzo Ferrante +4
cs.CVcs.LGarXiv:1703.02161v22017Learning End-to-End Lossy Image Compression: A Benchmark
Yueyu Hu, Wenhan Yang, Zhan Ma +1
eess.IVcs.CVarXiv:2002.03711v42020Memory-Efficient Incremental Learning Through Feature Adaptation
Ahmet Iscen, Jeffrey Zhang, Svetlana Lazebnik +1
cs.CVarXiv:2004.00713v220203D Object Reconstruction from a Single Depth View with Adversarial Learning
Bo Yang, Hongkai Wen, Sen Wang +3
cs.CVcs.AIcs.LGarXiv:1708.07969v12017Conditional Channel Gated Networks for Task-Aware Continual Learning
Davide Abati, Jakub Tomczak, Tijmen Blankevoort +3
cs.CVcs.LGstat.MLarXiv:2004.00070v12020MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning
Ke Wang, Houxing Ren, Aojun Zhou +7
cs.CLcs.AIcs.CVarXiv:2310.03731v12023MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Kaining Ying, Fanqing Meng, Jin Wang +19
cs.CVarXiv:2404.16006v12024Exploiting Image-trained CNN Architectures for Unconstrained Video Classification
Shengxin Zha, Florian Luisier, Walter Andrews +2
cs.CVarXiv:1503.04144v32015Re-initialization Free Level Set Evolution via Reaction Diffusion
Kaihua Zhang, Lei Zhang, Huihui Song +1
cs.CVarXiv:1112.1496v32011Learning Progressive Modality-shared Transformers for Effective Visible-Infrared Person Re-identification
Hu Lu, Xuezhang Zou, Pingping Zhang
cs.CVcs.IRcs.MMarXiv:2212.00226v12022VanillaNet: the Power of Minimalism in Deep Learning
Hanting Chen, Yunhe Wang, Jianyuan Guo +1
cs.CVarXiv:2305.12972v22023RenderDiffusion: Image Diffusion for 3D Reconstruction, Inpainting and Generation
Titas Anciukevičius, Zexiang Xu, Matthew Fisher +4
cs.CVcs.LGarXiv:2211.09869v42022Retrosynthesis of Synthetic Media for Explainable AI Provenance Forensics
Yijie Lin, Ching-Chun Chang, Isao Echizen +2
cs.CRcs.AIcs.CVarXiv:2609.02268v12026