Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
13,081 to 13,140 of 18,822
A Sufficient Condition for Convergences of Adam and RMSProp
Fangyu Zou, Li Shen, Zequn Jie +2
cs.LGcs.CVmath.NAarXiv:1811.09358v32018Simple diffusion: End-to-end diffusion for high resolution images
Emiel Hoogeboom, Jonathan Heek, Tim Salimans
cs.CVcs.LGstat.MLarXiv:2301.11093v22023Attention Branch Network: Learning of Attention Mechanism for Visual Explanation
Hiroshi Fukui, Tsubasa Hirakawa, Takayoshi Yamashita +1
cs.CVarXiv:1812.10025v22018Unsupervised Discovery of Interpretable Directions in the GAN Latent Space
Andrey Voynov, Artem Babenko
cs.LGcs.CVstat.MLarXiv:2002.03754v32020Revisiting Feature Prediction for Learning Visual Representations from Video
Adrien Bardes, Quentin Garrido, Jean Ponce +5
cs.CVcs.AIcs.LGarXiv:2404.08471v12024CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models
Yucheng Zhou, Peng Luo, Qianning Wang +2
cs.CLcs.CVarXiv:2608.26147v12026Socialized Detector Learning: Trajectory-Guided and Reciprocal Distillation for Heterogeneous Object Detectors
Weihao Li, Yunqi Zhu, Zhihe Fan +5
cs.CVarXiv:2608.25836v12026Training-Free VLM Personalization via Calibrated Residual Decoding
Jiaao Yu, Yujian Ma, Xianming Hu +2
cs.CVcs.AIarXiv:2608.22263v12026CertVLA: Certified Defense against Physical Visual Attacks for Vision-Language-Action Models
Hui Lu, Zhijie Peng, Yuqi Lin +8
cs.CVcs.AIarXiv:2608.20791v12026Data Efficient and Weakly Supervised Computational Pathology on Whole Slide Images
Ming Y. Lu, Drew F. K. Williamson, Tiffany Y. Chen +3
eess.IVcs.CVcs.LGarXiv:2004.09666v22020Deep Domain-Adversarial Image Generation for Domain Generalisation
Kaiyang Zhou, Yongxin Yang, Timothy Hospedales +1
cs.CVarXiv:2003.06054v12020FU-Mamba: A Frequency-Enhanced Dynamic Scanning Framework for Oralscan Image Segmentation
Xinxin Zhao, Jinpeng Ye, Bo Wei +8
cs.CVarXiv:2608.26607v12026R2RNet: Low-light Image Enhancement via Real-low to Real-normal Network
Jiang Hai, Zhu Xuan, Songchen Han +4
cs.CVeess.IVarXiv:2106.14501v22021Federated Learning for Medical Image Analysis: A Survey
Hao Guan, Pew-Thian Yap, Andrea Bozoki +1
cs.CVeess.IVarXiv:2306.05980v42023UPSNet: A Unified Panoptic Segmentation Network
Yuwen Xiong, Renjie Liao, Hengshuang Zhao +4
cs.CVarXiv:1901.03784v22019Using satellite imagery to understand and promote sustainable development
Marshall Burke, Anne Driscoll, David B. Lobell +1
cs.CYcs.CVcs.LGarXiv:2010.06988v12020ClassVision: AI-Powered Classroom Attendance System
Ankit Kumar Aggarwal, Veerabhadra Rao Marellapudi, Ovadia Sutton +1
cs.CYcs.AIcs.CVarXiv:2608.26173v12026GPT-Driver: Learning to Drive with GPT
Jiageng Mao, Yuxi Qian, Junjie Ye +2
cs.CVcs.AIcs.CLarXiv:2310.01415v32023Efficient Long-Range Attention Network for Image Super-resolution
Xindong Zhang, Hui Zeng, Shi Guo +1
cs.CVarXiv:2203.06697v12022Ring Forcing: Towards Precise Long-Term Memory for Autoregressive Video Diffusion
Bowen Xue, Brandon Y. Feng, Chenguo Lin +6
cs.CVarXiv:2608.26794v12026Camera Calibration Using Inaccurate and Asynchronous Discrete GPS Trajectory from Drones
R. Yang, Y. Bar-Shalom, H. A. J. Huang
eess.SYcs.CVarXiv:2608.26548v12026Glass Surface Detection Grounded in 3D Visual Geometry
Yiwei Lu, Ke Xu, Tao Yan +3
cs.CVarXiv:2608.26752v12026MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA
Haowen Gu, Gensheng Pei, Zeren Sun +4
cs.CVcs.AIarXiv:2608.26848v12026Local Aggregation for Unsupervised Learning of Visual Embeddings
Chengxu Zhuang, Alex Lin Zhai, Daniel Yamins
cs.CVcs.AIarXiv:1903.12355v22019ASF-YOLO: A Novel YOLO Model with Attentional Scale Sequence Fusion for Cell Instance Segmentation
Ming Kang, Chee-Ming Ting, Fung Fung Ting +1
cs.CVeess.SPstat.AParXiv:2312.06458v22023TrafficPredict: Trajectory Prediction for Heterogeneous Traffic-Agents
Yuexin Ma, Xinge Zhu, Sibo Zhang +3
cs.CVcs.ROarXiv:1811.02146v52018Decomposing NeRF for Editing via Feature Field Distillation
Sosuke Kobayashi, Eiichi Matsumoto, Vincent Sitzmann
cs.CVcs.GRarXiv:2205.15585v22022What makes fake images detectable? Understanding properties that generalize
Lucy Chai, David Bau, Ser-Nam Lim +1
cs.CVarXiv:2008.10588v12020Cyclical Annealing Schedule: A Simple Approach to Mitigating KL Vanishing
Hao Fu, Chunyuan Li, Xiaodong Liu +3
cs.LGcs.AIcs.CLarXiv:1903.10145v32019Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Chengyue Wu, Xiaokang Chen, Zhiyu Wu +8
cs.CVcs.AIcs.CLarXiv:2410.13848v12024Evaluate the Malignancy of Pulmonary Nodules Using the 3D Deep Leaky Noisy-or Network
Fangzhou Liao, Ming Liang, Zhe Li +2
cs.CVarXiv:1711.08324v12017Multi-Garment Net: Learning to Dress 3D People from Images
Bharat Lal Bhatnagar, Garvita Tiwari, Christian Theobalt +1
cs.CVarXiv:1908.06903v22019TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering
Yushi Hu, Benlin Liu, Jungo Kasai +4
cs.CVarXiv:2303.11897v32023Skip-GANomaly: Skip Connected and Adversarially Trained Encoder-Decoder Anomaly Detection
Samet Akçay, Amir Atapour-Abarghouei, Toby P. Breckon
cs.CVcs.LGarXiv:1901.08954v12019VITON-HD: High-Resolution Virtual Try-On via Misalignment-Aware Normalization
Seunghwan Choi, Sunghyun Park, Minsoo Lee +1
cs.CVarXiv:2103.16874v22021A Study and Comparison of Human and Deep Learning Recognition Performance Under Visual Distortions
Samuel Dodge, Lina Karam
cs.CVarXiv:1705.02498v12017Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Chunting Zhou, Lili Yu, Arun Babu +7
cs.AIcs.CVarXiv:2408.11039v12024Swin-UMamba: Mamba-based UNet with ImageNet-based pretraining
Jiarun Liu, Hao Yang, Hong-Yu Zhou +8
eess.IVcs.CVcs.LGarXiv:2402.03302v22024Attention-Aware Compositional Network for Person Re-identification
Jing Xu, Rui Zhao, Feng Zhu +2
cs.CVarXiv:1805.03344v22018Assessment of algorithms for mitosis detection in breast cancer histopathology images
Mitko Veta, Paul J. van Diest, Stefan M. Willems +26
cs.CVarXiv:1411.5825v12014Multi-Agent Tensor Fusion for Contextual Trajectory Prediction
Tianyang Zhao, Yifei Xu, Mathew Monfort +5
cs.CVcs.LGarXiv:1904.04776v22019FlowFormer: A Transformer Architecture for Optical Flow
Zhaoyang Huang, Xiaoyu Shi, Chao Zhang +5
cs.CVarXiv:2203.16194v42022FaceShifter: Towards High Fidelity And Occlusion Aware Face Swapping
Lingzhi Li, Jianmin Bao, Hao Yang +2
cs.CVarXiv:1912.13457v32019Depth Pro: Sharp Monocular Metric Depth in Less Than a Second
Aleksei Bochkovskii, Amaël Delaunoy, Hugo Germain +4
cs.CVcs.LGarXiv:2410.02073v22024Summaries:한국어Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual Representation
Hang Zhou, Yasheng Sun, Wayne Wu +3
cs.CVcs.LGcs.MMarXiv:2104.11116v12021Kornia: an Open Source Differentiable Computer Vision Library for PyTorch
Edgar Riba, Dmytro Mishkin, Daniel Ponsa +2
cs.CVarXiv:1910.02190v22019Domain Adaptive Neural Networks for Object Recognition
Muhammad Ghifary, W. Bastiaan Kleijn, Mengjie Zhang
cs.CVcs.AIcs.LGarXiv:1409.6041v12014Deep Metric Learning with Hierarchical Triplet Loss
Weifeng Ge, Weilin Huang, Dengke Dong +1
cs.CVarXiv:1810.06951v12018Stacked Generative Adversarial Networks
Xun Huang, Yixuan Li, Omid Poursaeed +2
cs.CVcs.LGcs.NEarXiv:1612.04357v42016Separable Self-attention for Mobile Vision Transformers
Sachin Mehta, Mohammad Rastegari
cs.CVcs.AIcs.LGarXiv:2206.02680v12022Data Augmentation by Pairing Samples for Images Classification
Hiroshi Inoue
cs.LGcs.CVstat.MLarXiv:1801.02929v22018Deformable Part Models are Convolutional Neural Networks
Ross Girshick, Forrest Iandola, Trevor Darrell +1
cs.CVarXiv:1409.5403v22014TransMed: Transformers Advance Multi-modal Medical Image Classification
Yin Dai, Yifan Gao
cs.CVarXiv:2103.05940v12021SuperYOLO: Super Resolution Assisted Object Detection in Multimodal Remote Sensing Imagery
Jiaqing Zhang, Jie Lei, Weiying Xie +3
cs.CVarXiv:2209.13351v22022MMRotate: A Rotated Object Detection Benchmark using PyTorch
Yue Zhou, Xue Yang, Gefan Zhang +9
cs.CVcs.AIarXiv:2204.13317v42022LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
Bin Zhu, Bin Lin, Munan Ning +11
cs.CVcs.AIarXiv:2310.01852v72023PatchmatchNet: Learned Multi-View Patchmatch Stereo
Fangjinhua Wang, Silvano Galliani, Christoph Vogel +2
cs.CVarXiv:2012.01411v12020Deep MANTA: A Coarse-to-fine Many-Task Network for joint 2D and 3D vehicle analysis from monocular image
Florian Chabot, Mohamed Chaouch, Jaonary Rabarisoa +2
cs.CVarXiv:1703.07570v12017Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language Model
Yu Du, Fangyun Wei, Zihe Zhang +3
cs.CVarXiv:2203.14940v12022CAT3D: Create Anything in 3D with Multi-View Diffusion Models
Ruiqi Gao, Aleksander Holynski, Philipp Henzler +5
cs.CVarXiv:2405.10314v12024