Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
14,701 to 14,760 of 18,866
Few-Shot Unsupervised Image-to-Image Translation
Ming-Yu Liu, Xun Huang, Arun Mallya +4
cs.CVcs.AIcs.GRarXiv:1905.01723v22019TernausNet: U-Net with VGG11 Encoder Pre-Trained on ImageNet for Image Segmentation
Vladimir Iglovikov, Alexey Shvets
cs.CVarXiv:1801.05746v12018CFLOW-AD: Real-Time Unsupervised Anomaly Detection with Localization via Conditional Normalizing Flows
Denis Gudovskiy, Shun Ishizaka, Kazuki Kozuka
cs.CVcs.AIcs.LGarXiv:2107.12571v12021SplaTAM: Splat, Track & Map 3D Gaussians for Dense RGB-D SLAM
Nikhil Keetha, Jay Karhade, Krishna Murthy Jatavallabhula +4
cs.CVcs.AIcs.ROarXiv:2312.02126v32023ORBITALIF: An Efficient Spiking Federated Learning Framework for Onboard Cloud Removal
Bohan Zhang, Chenyu Xu, Yijie Mao +1
cs.NEcs.AIcs.CVarXiv:2608.24073v12026Contextual Transformer Networks for Visual Recognition
Yehao Li, Ting Yao, Yingwei Pan +1
cs.CVcs.AIcs.LGarXiv:2107.12292v12021IterCAD: Iterative Program Repair for CAD Code Generation from Orthographic Views
Yuchuan Wu, Ke Niu, Haiyang Yu +3
cs.CVcs.AIarXiv:2608.24020v12026Adapting Vision-Language Models for E-commerce Understanding at Scale
Matteo Nulli, Vladimir Orshulevich, Tala Bazazo +9
cs.CVcs.AIarXiv:2602.11733v12026DeepSight: An All-in-One LM Safety Toolkit
Bo Zhang, Jiaxuan Guo, Lijun Li +17
cs.CLcs.AIcs.CRarXiv:2602.12092v12026VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
Haoxin Chen, Yong Zhang, Xiaodong Cun +4
cs.CVarXiv:2401.09047v12024DFANet: Deep Feature Aggregation for Real-Time Semantic Segmentation
Hanchao Li, Pengfei Xiong, Haoqiang Fan +1
cs.CVarXiv:1904.02216v12019CellMaster: Collaborative Cell Type Annotation in Single-Cell Analysis
Zhen Wang, Yiming Gao, Jieyuan Liu +10
q-bio.GNcs.AIcs.CVarXiv:2602.13346v12026UniWeTok: An Unified Binary Tokenizer with Codebook Size $\mathit{2^{128}}$ for Unified Multimodal Large Language Model
Shaobin Zhuang, Yuang Ai, Jiaming Han +12
cs.CVcs.AIarXiv:2602.14178v32026Improving the Robustness of Deep Neural Networks via Stability Training
Stephan Zheng, Yang Song, Thomas Leung +1
cs.CVcs.LGarXiv:1604.04326v12016Medical Image Synthesis for Data Augmentation and Anonymization using Generative Adversarial Networks
Hoo-Chang Shin, Neil A Tenenholtz, Jameson K Rogers +5
cs.CVcs.LGstat.MLarXiv:1807.10225v22018Exploiting Diffusion Prior for Real-World Image Super-Resolution
Jianyi Wang, Zongsheng Yue, Shangchen Zhou +2
cs.CVarXiv:2305.07015v42023Real Image Denoising with Feature Attention
Saeed Anwar, Nick Barnes
cs.CVcs.LGarXiv:1904.07396v22019Focusing Where Vision Matters: Selective Training for Large Vision Language Models via Visual Information Gain
Seulbi Lee, Sangheum Hwang
cs.CVarXiv:2602.17186v22026Deep Learning Super Resolution for Satellite Cloud Mask Downscaling
Angelos Georgakis, Valentina Kanaki, Giorgos Giannopoulos +4
cs.CVcs.AIarXiv:2608.24715v12026CANet: Class-Agnostic Segmentation Networks with Iterative Refinement and Attentive Few-Shot Learning
Chi Zhang, Guosheng Lin, Fayao Liu +2
cs.CVarXiv:1903.02351v12019EMFE: A lightweight, explainable machine learning framework for malaria cell classification
Md Abdullah Al Kafi, Walayat Hussain, Mousumi Karmakar +2
cs.CVarXiv:2608.24793v12026MoRL: Reinforced Reasoning for Unified Motion Understanding and Generation
Hongpeng Wang, Zeyu Zhang, Wenhao Li +1
cs.CVarXiv:2602.14534v12026Learning with a Wasserstein Loss
Charlie Frogner, Chiyuan Zhang, Hossein Mobahi +2
cs.LGcs.CVstat.MLarXiv:1506.05439v32015PointASNL: Robust Point Clouds Processing using Nonlocal Neural Networks with Adaptive Sampling
Xu Yan, Chaoda Zheng, Zhen Li +2
cs.CVarXiv:2003.00492v32020Visual Persuasion: What Influences Decisions of Vision-Language Models?
Manuel Cherep, Pranav M R, Pattie Maes +1
cs.CVcs.AIarXiv:2602.15278v22026HERO: Learning Humanoid End-Effector Control for Visual Whole-Body Open-Vocabulary Object Grasping
Runpei Dong, Ziyan Li, Arjun Gupta +2
cs.ROcs.CVarXiv:2602.16705v32026StereoAdapter-2: Globally Structure-Consistent Underwater Stereo Depth Estimation
Zeyu Ren, Xiang Li, Yiran Wang +2
cs.CVarXiv:2602.16915v12026V2VNet: Vehicle-to-Vehicle Communication for Joint Perception and Prediction
Tsun-Hsuan Wang, Sivabalan Manivasagam, Ming Liang +4
cs.CVarXiv:2008.07519v12020MoTE: Mixture of Task Experts for Multi-Task Video Understanding
Muhammad Asad Ali, Umar Khan, Nadia Robertini +1
cs.CVcs.LGarXiv:2608.24763v12026Retrieval-Augmented Generation for AI-Generated Content: A Survey
Penghao Zhao, Hailin Zhang, Qinhan Yu +7
cs.CVarXiv:2402.19473v62024Deep MR to CT Synthesis using Unpaired Data
Jelmer M. Wolterink, Anna M. Dinkla, Mark H. F. Savenije +3
cs.CVarXiv:1708.01155v12017Action Recognition using Visual Attention
Shikhar Sharma, Ryan Kiros, Ruslan Salakhutdinov
cs.LGcs.CVarXiv:1511.04119v32015GlanceWAM: Sparse Test-Time Imagination for World-Action Models
Linhan Wang, Zijian An, Mingyuan Zhang +7
cs.CVarXiv:2608.23927v12026YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark
Ning Xu, Linjie Yang, Yuchen Fan +4
cs.CVcs.AIarXiv:1809.03327v12018Summaries:한국어CARE: Camera-Residual Reserves for First Sightings in Adaptive LiDAR Sensing
Jiachen Gong, Yun Li, Ehsan Javanmardi +2
cs.CVcs.ROarXiv:2608.24282v12026SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering
Bo Liu, Li-Ming Zhan, Li Xu +3
cs.CVcs.AIcs.CLarXiv:2102.09542v12021Deep Learning in Medical Image Registration: A Survey
Grant Haskins, Uwe Kruger, Pingkun Yan
q-bio.QMcs.CVeess.IVarXiv:1903.02026v22019Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
Lijun Yu, José Lezama, Nitesh B. Gundavarapu +13
cs.CVcs.AIcs.MMarXiv:2310.05737v32023Cascaded Recurrent Neural Networks for Hyperspectral Image Classification
Renlong Hang, Qingshan Liu, Danfeng Hong +1
cs.CVarXiv:1902.10858v12019DeepFuse: A Deep Unsupervised Approach for Exposure Fusion with Extreme Exposure Image Pairs
K. Ram Prabhakar, V. Sai Srikar, R. Venkatesh Babu
cs.CVarXiv:1712.07384v12017DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
Aditya Kumar Singh, Hitesh Kandala, Pratik Prabhanjan Brahma +2
cs.CVcs.AIarXiv:2602.18846v22026Hepato-LLaVA: An Expert MLLM with Sparse Topo-Pack Attention for Hepatocellular Pathology Analysis on Whole Slide Images
Yuxuan Yang, Zhonghao Yan, Yi Zhang +6
cs.CVarXiv:2602.19424v32026Underwater Single Image Color Restoration Using Haze-Lines and a New Quantitative Dataset
Dana Berman, Deborah Levy, Shai Avidan +1
cs.CVarXiv:1811.01343v32018Communication-Inspired Tokenization for Structured Image Representations
Aram Davtyan, Yusuf Sahin, Yasaman Haghighi +4
cs.CVcs.AIcs.LGarXiv:2602.20731v12026AVEC 2016 - Depression, Mood, and Emotion Recognition Workshop and Challenge
Michel Valstar, Jonathan Gratch, Bjorn Schuller +7
cs.CVcs.HCcs.MMarXiv:1605.01600v42016How to Take a Memorable Picture? Empowering Users with Actionable Feedback
Francesco Laiti, Davide Talon, Jacopo Staiano +1
cs.CVarXiv:2602.21877v22026DER: Dynamically Expandable Representation for Class Incremental Learning
Shipeng Yan, Jiangwei Xie, Xuming He
cs.CVcs.LGarXiv:2103.16788v12021Half-Truths Break Similarity-Based Retrieval
Bora Kargi, Arnas Uselis, Seong Joon Oh
cs.CVarXiv:2602.23906v12026MicroVerse: A Preliminary Exploration Toward a Micro-World Simulation
Rongsheng Wang, Minghao Wu, Hongru Zhou +4
cs.AIcs.CVarXiv:2603.00585v12026Spatial Attentive Single-Image Deraining with a High Quality Real Rain Dataset
Tianyu Wang, Xin Yang, Ke Xu +3
cs.CVarXiv:1904.01538v22019X-MULTI: VLM-based Imaging Factor Disentanglement for Factor-Aware Image Synthesis
Sonali Godavarthy, Matthias Neuwirth-Trapp, Tim-Felix Faasch +4
cs.CVcs.ROarXiv:2608.24563v12026SegMamba: Long-range Sequential Modeling Mamba For 3D Medical Image Segmentation
Zhaohu Xing, Tian Ye, Yijun Yang +2
cs.CVarXiv:2401.13560v42024From Blind Spots to Gains: Diagnostic-Driven Iterative Training for Large Multimodal Models
Hongrui Jia, Chaoya Jiang, Yongrui Heng +2
cs.CVarXiv:2602.22859v22026MobileNetV4 -- Universal Models for the Mobile Ecosystem
Danfeng Qin, Chas Leichner, Manolis Delakis +11
cs.CVarXiv:2404.10518v22024Retrieve and Segment: Are a Few Examples Enough to Bridge the Supervision Gap in Open-Vocabulary Segmentation?
Tilemachos Aravanis, Vladan Stojnić, Bill Psomas +2
cs.CVarXiv:2602.23339v12026Faster Segment Anything: Towards Lightweight SAM for Mobile Applications
Chaoning Zhang, Dongshen Han, Yu Qiao +4
cs.CVarXiv:2306.14289v22023DLEBench: Evaluating Small-scale Object Editing Ability for Instruction-based Image Editing Model
Shibo Hong, Boxian Ai, Jun Kuang +5
cs.CVcs.AIarXiv:2602.23622v22026Enhancing Spatial Understanding in Image Generation via Reward Modeling
Zhenyu Tang, Chaoran Feng, Yufan Deng +5
cs.CVarXiv:2602.24233v12026Towards Reliable AI-Based Histological Staining: A Systematic Study of Scaling and Uncertainty in Unpaired Generative Models
Qasim Siddiqui, Adrian Friebel, Maiju Myllys +4
cs.CVarXiv:2608.24626v12026Variational Adversarial Active Learning
Samarth Sinha, Sayna Ebrahimi, Trevor Darrell
cs.LGcs.CVstat.MLarXiv:1904.00370v32019