Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

13,081 to 13,140 of 18,822

  1. A Sufficient Condition for Convergences of Adam and RMSProp

    Fangyu Zou, Li Shen, Zequn Jie +2

    cs.LGcs.CVmath.NAarXiv:1811.09358v32018
  2. Simple diffusion: End-to-end diffusion for high resolution images

    Emiel Hoogeboom, Jonathan Heek, Tim Salimans

    cs.CVcs.LGstat.MLarXiv:2301.11093v22023
  3. Attention Branch Network: Learning of Attention Mechanism for Visual Explanation

    Hiroshi Fukui, Tsubasa Hirakawa, Takayoshi Yamashita +1

    cs.CVarXiv:1812.10025v22018
  4. Unsupervised Discovery of Interpretable Directions in the GAN Latent Space

    Andrey Voynov, Artem Babenko

    cs.LGcs.CVstat.MLarXiv:2002.03754v32020
  5. Revisiting Feature Prediction for Learning Visual Representations from Video

    Adrien Bardes, Quentin Garrido, Jean Ponce +5

    cs.CVcs.AIcs.LGarXiv:2404.08471v12024
  6. CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models

    Yucheng Zhou, Peng Luo, Qianning Wang +2

    cs.CLcs.CVarXiv:2608.26147v12026
  7. Socialized Detector Learning: Trajectory-Guided and Reciprocal Distillation for Heterogeneous Object Detectors

    Weihao Li, Yunqi Zhu, Zhihe Fan +5

    cs.CVarXiv:2608.25836v12026
  8. Training-Free VLM Personalization via Calibrated Residual Decoding

    Jiaao Yu, Yujian Ma, Xianming Hu +2

    cs.CVcs.AIarXiv:2608.22263v12026
  9. CertVLA: Certified Defense against Physical Visual Attacks for Vision-Language-Action Models

    Hui Lu, Zhijie Peng, Yuqi Lin +8

    cs.CVcs.AIarXiv:2608.20791v12026
  10. Data Efficient and Weakly Supervised Computational Pathology on Whole Slide Images

    Ming Y. Lu, Drew F. K. Williamson, Tiffany Y. Chen +3

    eess.IVcs.CVcs.LGarXiv:2004.09666v22020
  11. Deep Domain-Adversarial Image Generation for Domain Generalisation

    Kaiyang Zhou, Yongxin Yang, Timothy Hospedales +1

    cs.CVarXiv:2003.06054v12020
  12. FU-Mamba: A Frequency-Enhanced Dynamic Scanning Framework for Oralscan Image Segmentation

    Xinxin Zhao, Jinpeng Ye, Bo Wei +8

    cs.CVarXiv:2608.26607v12026
  13. R2RNet: Low-light Image Enhancement via Real-low to Real-normal Network

    Jiang Hai, Zhu Xuan, Songchen Han +4

    cs.CVeess.IVarXiv:2106.14501v22021
  14. Federated Learning for Medical Image Analysis: A Survey

    Hao Guan, Pew-Thian Yap, Andrea Bozoki +1

    cs.CVeess.IVarXiv:2306.05980v42023
  15. UPSNet: A Unified Panoptic Segmentation Network

    Yuwen Xiong, Renjie Liao, Hengshuang Zhao +4

    cs.CVarXiv:1901.03784v22019
  16. Using satellite imagery to understand and promote sustainable development

    Marshall Burke, Anne Driscoll, David B. Lobell +1

    cs.CYcs.CVcs.LGarXiv:2010.06988v12020
  17. ClassVision: AI-Powered Classroom Attendance System

    Ankit Kumar Aggarwal, Veerabhadra Rao Marellapudi, Ovadia Sutton +1

    cs.CYcs.AIcs.CVarXiv:2608.26173v12026
  18. GPT-Driver: Learning to Drive with GPT

    Jiageng Mao, Yuxi Qian, Junjie Ye +2

    cs.CVcs.AIcs.CLarXiv:2310.01415v32023
  19. Efficient Long-Range Attention Network for Image Super-resolution

    Xindong Zhang, Hui Zeng, Shi Guo +1

    cs.CVarXiv:2203.06697v12022
  20. Ring Forcing: Towards Precise Long-Term Memory for Autoregressive Video Diffusion

    Bowen Xue, Brandon Y. Feng, Chenguo Lin +6

    cs.CVarXiv:2608.26794v12026
  21. Camera Calibration Using Inaccurate and Asynchronous Discrete GPS Trajectory from Drones

    R. Yang, Y. Bar-Shalom, H. A. J. Huang

    eess.SYcs.CVarXiv:2608.26548v12026
  22. Glass Surface Detection Grounded in 3D Visual Geometry

    Yiwei Lu, Ke Xu, Tao Yan +3

    cs.CVarXiv:2608.26752v12026
  23. MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA

    Haowen Gu, Gensheng Pei, Zeren Sun +4

    cs.CVcs.AIarXiv:2608.26848v12026
  24. Local Aggregation for Unsupervised Learning of Visual Embeddings

    Chengxu Zhuang, Alex Lin Zhai, Daniel Yamins

    cs.CVcs.AIarXiv:1903.12355v22019
  25. ASF-YOLO: A Novel YOLO Model with Attentional Scale Sequence Fusion for Cell Instance Segmentation

    Ming Kang, Chee-Ming Ting, Fung Fung Ting +1

    cs.CVeess.SPstat.AParXiv:2312.06458v22023
  26. TrafficPredict: Trajectory Prediction for Heterogeneous Traffic-Agents

    Yuexin Ma, Xinge Zhu, Sibo Zhang +3

    cs.CVcs.ROarXiv:1811.02146v52018
  27. Decomposing NeRF for Editing via Feature Field Distillation

    Sosuke Kobayashi, Eiichi Matsumoto, Vincent Sitzmann

    cs.CVcs.GRarXiv:2205.15585v22022
  28. What makes fake images detectable? Understanding properties that generalize

    Lucy Chai, David Bau, Ser-Nam Lim +1

    cs.CVarXiv:2008.10588v12020
  29. Cyclical Annealing Schedule: A Simple Approach to Mitigating KL Vanishing

    Hao Fu, Chunyuan Li, Xiaodong Liu +3

    cs.LGcs.AIcs.CLarXiv:1903.10145v32019
  30. Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

    Chengyue Wu, Xiaokang Chen, Zhiyu Wu +8

    cs.CVcs.AIcs.CLarXiv:2410.13848v12024
  31. Evaluate the Malignancy of Pulmonary Nodules Using the 3D Deep Leaky Noisy-or Network

    Fangzhou Liao, Ming Liang, Zhe Li +2

    cs.CVarXiv:1711.08324v12017
  32. Multi-Garment Net: Learning to Dress 3D People from Images

    Bharat Lal Bhatnagar, Garvita Tiwari, Christian Theobalt +1

    cs.CVarXiv:1908.06903v22019
  33. TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

    Yushi Hu, Benlin Liu, Jungo Kasai +4

    cs.CVarXiv:2303.11897v32023
  34. Skip-GANomaly: Skip Connected and Adversarially Trained Encoder-Decoder Anomaly Detection

    Samet Akçay, Amir Atapour-Abarghouei, Toby P. Breckon

    cs.CVcs.LGarXiv:1901.08954v12019
  35. VITON-HD: High-Resolution Virtual Try-On via Misalignment-Aware Normalization

    Seunghwan Choi, Sunghyun Park, Minsoo Lee +1

    cs.CVarXiv:2103.16874v22021
  36. A Study and Comparison of Human and Deep Learning Recognition Performance Under Visual Distortions

    Samuel Dodge, Lina Karam

    cs.CVarXiv:1705.02498v12017
  37. Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

    Chunting Zhou, Lili Yu, Arun Babu +7

    cs.AIcs.CVarXiv:2408.11039v12024
  38. Swin-UMamba: Mamba-based UNet with ImageNet-based pretraining

    Jiarun Liu, Hao Yang, Hong-Yu Zhou +8

    eess.IVcs.CVcs.LGarXiv:2402.03302v22024
  39. Attention-Aware Compositional Network for Person Re-identification

    Jing Xu, Rui Zhao, Feng Zhu +2

    cs.CVarXiv:1805.03344v22018
  40. Assessment of algorithms for mitosis detection in breast cancer histopathology images

    Mitko Veta, Paul J. van Diest, Stefan M. Willems +26

    cs.CVarXiv:1411.5825v12014
  41. Multi-Agent Tensor Fusion for Contextual Trajectory Prediction

    Tianyang Zhao, Yifei Xu, Mathew Monfort +5

    cs.CVcs.LGarXiv:1904.04776v22019
  42. FlowFormer: A Transformer Architecture for Optical Flow

    Zhaoyang Huang, Xiaoyu Shi, Chao Zhang +5

    cs.CVarXiv:2203.16194v42022
  43. FaceShifter: Towards High Fidelity And Occlusion Aware Face Swapping

    Lingzhi Li, Jianmin Bao, Hao Yang +2

    cs.CVarXiv:1912.13457v32019
  44. Depth Pro: Sharp Monocular Metric Depth in Less Than a Second

    Aleksei Bochkovskii, Amaël Delaunoy, Hugo Germain +4

    cs.CVcs.LGarXiv:2410.02073v22024
    Summaries:한국어
  45. Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual Representation

    Hang Zhou, Yasheng Sun, Wayne Wu +3

    cs.CVcs.LGcs.MMarXiv:2104.11116v12021
  46. Kornia: an Open Source Differentiable Computer Vision Library for PyTorch

    Edgar Riba, Dmytro Mishkin, Daniel Ponsa +2

    cs.CVarXiv:1910.02190v22019
  47. Domain Adaptive Neural Networks for Object Recognition

    Muhammad Ghifary, W. Bastiaan Kleijn, Mengjie Zhang

    cs.CVcs.AIcs.LGarXiv:1409.6041v12014
  48. Deep Metric Learning with Hierarchical Triplet Loss

    Weifeng Ge, Weilin Huang, Dengke Dong +1

    cs.CVarXiv:1810.06951v12018
  49. Stacked Generative Adversarial Networks

    Xun Huang, Yixuan Li, Omid Poursaeed +2

    cs.CVcs.LGcs.NEarXiv:1612.04357v42016
  50. Separable Self-attention for Mobile Vision Transformers

    Sachin Mehta, Mohammad Rastegari

    cs.CVcs.AIcs.LGarXiv:2206.02680v12022
  51. Data Augmentation by Pairing Samples for Images Classification

    Hiroshi Inoue

    cs.LGcs.CVstat.MLarXiv:1801.02929v22018
  52. Deformable Part Models are Convolutional Neural Networks

    Ross Girshick, Forrest Iandola, Trevor Darrell +1

    cs.CVarXiv:1409.5403v22014
  53. TransMed: Transformers Advance Multi-modal Medical Image Classification

    Yin Dai, Yifan Gao

    cs.CVarXiv:2103.05940v12021
  54. SuperYOLO: Super Resolution Assisted Object Detection in Multimodal Remote Sensing Imagery

    Jiaqing Zhang, Jie Lei, Weiying Xie +3

    cs.CVarXiv:2209.13351v22022
  55. MMRotate: A Rotated Object Detection Benchmark using PyTorch

    Yue Zhou, Xue Yang, Gefan Zhang +9

    cs.CVcs.AIarXiv:2204.13317v42022
  56. LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

    Bin Zhu, Bin Lin, Munan Ning +11

    cs.CVcs.AIarXiv:2310.01852v72023
  57. PatchmatchNet: Learned Multi-View Patchmatch Stereo

    Fangjinhua Wang, Silvano Galliani, Christoph Vogel +2

    cs.CVarXiv:2012.01411v12020
  58. Deep MANTA: A Coarse-to-fine Many-Task Network for joint 2D and 3D vehicle analysis from monocular image

    Florian Chabot, Mohamed Chaouch, Jaonary Rabarisoa +2

    cs.CVarXiv:1703.07570v12017
  59. Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language Model

    Yu Du, Fangyun Wei, Zihe Zhang +3

    cs.CVarXiv:2203.14940v12022
  60. CAT3D: Create Anything in 3D with Multi-View Diffusion Models

    Ruiqi Gao, Aleksander Holynski, Philipp Henzler +5

    cs.CVarXiv:2405.10314v12024