Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

2,281 to 2,340 of 18,801

  1. Side-Aware Boundary Localization for More Precise Object Detection

    Jiaqi Wang, Wenwei Zhang, Yuhang Cao +6

    cs.CVarXiv:1912.04260v22019
  2. Retinal Vessel Segmentation in Fundoscopic Images with Generative Adversarial Networks

    Jaemin Son, Sang Jun Park, Kyu-Hwan Jung

    cs.CVcs.LGarXiv:1706.09318v12017
  3. AirAnchor: Bridging Local and Global Spatial Information for Zero-Shot Aerial Vision-and-Language Navigation

    Shanwei Fan, Bin Zhang, Zhiwei Xu +4

    cs.CVcs.AIcs.ROarXiv:2609.08442v12026
  4. Document Understanding Dataset and Evaluation (DUDE)

    Jordy Van Landeghem, Rubén Tito, Łukasz Borchmann +10

    cs.CVcs.CLcs.LGarXiv:2305.08455v32023
  5. PSTNet: Point Spatio-Temporal Convolution on Point Cloud Sequences

    Hehe Fan, Xin Yu, Yuhang Ding +2

    cs.CVarXiv:2205.13713v12022
  6. Bootstrapping Semantic Segmentation with Regional Contrast

    Shikun Liu, Shuaifeng Zhi, Edward Johns +1

    cs.CVcs.LGarXiv:2104.04465v42021
  7. 3DSS-Mamba: 3D-Spectral-Spatial Mamba for Hyperspectral Image Classification

    Yan He, Bing Tu, Bo Liu +2

    cs.CVeess.IVarXiv:2405.12487v22024
  8. Generalist Vision Foundation Models for Medical Imaging: A Case Study of Segment Anything Model on Zero-Shot Medical Segmentation

    Peilun Shi, Jianing Qiu, Sai Mu Dalike Abaxi +3

    cs.CVarXiv:2304.12637v22023
  9. Unsupervised Learning of Object Structure and Dynamics from Videos

    Matthias Minderer, Chen Sun, Ruben Villegas +3

    cs.CVarXiv:1906.07889v32019
  10. AEROBLADE: Training-Free Detection of Latent Diffusion Images Using Autoencoder Reconstruction Error

    Jonas Ricker, Denis Lukovnikov, Asja Fischer

    cs.CVarXiv:2401.17879v22024
  11. Goal-Oriented Gaze Estimation for Zero-Shot Learning

    Yang Liu, Lei Zhou, Xiao Bai +4

    cs.CVarXiv:2103.03433v12021
  12. ERNIE-ViLG 2.0: Improving Text-to-Image Diffusion Model with Knowledge-Enhanced Mixture-of-Denoising-Experts

    Zhida Feng, Zhenyu Zhang, Xintong Yu +12

    cs.CVcs.AIarXiv:2210.15257v22022
  13. FEANet: Feature-Enhanced Attention Network for RGB-Thermal Real-time Semantic Segmentation

    Fuqin Deng, Hua Feng, Mingjian Liang +7

    cs.CVarXiv:2110.08988v12021
  14. DifFace: Blind Face Restoration with Diffused Error Contraction

    Zongsheng Yue, Chen Change Loy

    cs.CVarXiv:2212.06512v42022
  15. Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-Resolution

    Shangchen Zhou, Peiqing Yang, Jianyi Wang +2

    cs.CVarXiv:2312.06640v12023
  16. Image-to-image Translation via Hierarchical Style Disentanglement

    Xinyang Li, Shengchuan Zhang, Jie Hu +6

    cs.CVarXiv:2103.01456v12021
  17. TokenCut: Segmenting Objects in Images and Videos with Self-supervised Transformer and Normalized Cut

    Yangtao Wang, Xi Shen, Yuan Yuan +5

    cs.CVstat.MLarXiv:2209.00383v32022
  18. Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation

    Fanqing Meng, Jiaqi Liao, Xinyu Tan +7

    cs.CVarXiv:2410.05363v12024
  19. Pose-Robust Face Recognition via Deep Residual Equivariant Mapping

    Kaidi Cao, Yu Rong, Cheng Li +2

    cs.CVarXiv:1803.00839v12018
  20. Tarsier: Recipes for Training and Evaluating Large Video Description Models

    Jiawei Wang, Liping Yuan, Yuchen Zhang +1

    cs.CVcs.LGarXiv:2407.00634v22024
  21. A Cloud-Based Hybrid Model for Real-Time Detection of BRTA-Approved Licence Plates Using YOLO Tiny and Haar Cascade

    Debashis Kar Suvra, Tahsina Farah Sanam

    cs.CVarXiv:2609.06507v12026
  22. Semantic Human Matting

    Quan Chen, Tiezheng Ge, Yanyu Xu +3

    cs.CVcs.GRcs.LGarXiv:1809.01354v22018
  23. Stochastic reconstruction of an oolitic limestone by generative adversarial networks

    Lukas Mosser, Olivier Dubrule, Martin J. Blunt

    cs.CVphysics.geo-phstat.MLarXiv:1712.02854v12017
  24. FPicker: Topology-Guided Evolution for Filament Tracing in Low-SNR Microscopy

    Tingyin Zhao, Mingtao Huang, Yuan Shen

    cs.CVcs.AIarXiv:2609.08305v12026
  25. Visual Imitation Made Easy

    Sarah Young, Dhiraj Gandhi, Shubham Tulsiani +3

    cs.ROcs.CVcs.LGarXiv:2008.04899v12020
  26. Vision-based Navigation with Language-based Assistance via Imitation Learning with Indirect Intervention

    Khanh Nguyen, Debadeepta Dey, Chris Brockett +1

    cs.LGcs.CLcs.CVarXiv:1812.04155v42018
  27. A Spectral-Spatial-Dependent Global Learning Framework for Insufficient and Imbalanced Hyperspectral Image Classification

    Qiqi Zhu, Weihuan Deng, Zhuo Zheng +5

    cs.CVarXiv:2105.14327v12021
  28. CALIPER: Clean Scenes Cannot Rank Physical Inference in Pretrained Visual Representations

    Aman Mehta, Riya Baviskar

    cs.ROcs.AIcs.CVarXiv:2609.08250v12026
  29. RenderFormer-V2: Neural Rendering with Heterogeneous Scene Primitives

    Chong Zeng, Yue Dong, Pieter Peers +2

    cs.CVcs.GRcs.LGarXiv:2609.05738v12026
  30. Fine-grained Recognition in the Wild: A Multi-Task Domain Adaptation Approach

    Timnit Gebru, Judy Hoffman, Li Fei-Fei

    cs.CVarXiv:1709.02476v12017
  31. Anatomical Priors in Convolutional Networks for Unsupervised Biomedical Segmentation

    Adrian V. Dalca, John Guttag, Mert R. Sabuncu

    cs.CVarXiv:1903.03148v12019
  32. Fusion-Mamba for Cross-modality Object Detection

    Wenhao Dong, Haodong Zhu, Shaohui Lin +6

    cs.CVcs.AIarXiv:2404.09146v12024
  33. CS-CLIP: Compositional Scene Graph-guided CLIP for Robust Compositional Reasoning

    SeongJun Jeong, Minjoon Jung, Woo Suk Choi +2

    cs.CVcs.AIarXiv:2609.08242v12026
  34. Data augmentation using synthetic data for time series classification with deep residual networks

    Hassan Ismail Fawaz, Germain Forestier, Jonathan Weber +2

    cs.CVcs.AIarXiv:1808.02455v12018
  35. Video Generative Adversarial Networks: A Review

    Nuha Aldausari, Arcot Sowmya, Nadine Marcus +1

    cs.CVcs.LGeess.IVarXiv:2011.02250v12020
  36. Monocular Object Instance Segmentation and Depth Ordering with CNNs

    Ziyu Zhang, Alexander G. Schwing, Sanja Fidler +1

    cs.CVarXiv:1505.03159v22015
  37. CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation

    Dejia Xu, Weili Nie, Chao Liu +4

    cs.CVarXiv:2406.02509v12024
  38. WSPolypNet: Weakly Supervised Polyp Localization in Colonoscopy Videos

    Giseong Hwang, Minjae Jo, Yeonghyeon Park +9

    cs.CVcs.AIarXiv:2609.08182v12026
  39. Exploring Disentangled Feature Representation Beyond Face Identification

    Yu Liu, Fangyin Wei, Jing Shao +3

    cs.CVcs.AIarXiv:1804.03487v12018
  40. PDE-GCN: Novel Architectures for Graph Neural Networks Motivated by Partial Differential Equations

    Moshe Eliasof, Eldad Haber, Eran Treister

    cs.LGcs.CVcs.NEarXiv:2108.01938v22021
  41. PyramidCLIP: Hierarchical Feature Alignment for Vision-language Model Pretraining

    Yuting Gao, Jinfeng Liu, Zihan Xu +4

    cs.CVcs.AIarXiv:2204.14095v22022
  42. Score-Based Diffusion Models as Principled Priors for Inverse Imaging

    Berthy T. Feng, Jamie Smith, Michael Rubinstein +3

    cs.CVarXiv:2304.11751v22023
  43. MotionTrack: Learning Robust Short-term and Long-term Motions for Multi-Object Tracking

    Zheng Qin, Sanping Zhou, Le Wang +3

    cs.CVarXiv:2303.10404v22023
  44. RMDL: Recalibrated multi-instance deep learning for whole slide gastric image classification

    Shujun Wang, Yaxi Zhu, Lequan Yu +5

    cs.CVarXiv:2010.06440v12020
  45. Two Causal Principles for Improving Visual Dialog

    Jiaxin Qi, Yulei Niu, Jianqiang Huang +1

    cs.CVcs.CLarXiv:1911.10496v32019
  46. YOLOv1 to YOLOv10: A comprehensive review of YOLO variants and their application in the agricultural domain

    Mujadded Al Rabbani Alif, Muhammad Hussain

    cs.CVarXiv:2406.10139v12024
  47. Deep Texture-Aware Features for Camouflaged Object Detection

    Jingjing Ren, Xiaowei Hu, Lei Zhu +5

    cs.CVarXiv:2102.02996v12021
  48. Structured Scene Memory for Vision-Language Navigation

    Hanqing Wang, Wenguan Wang, Wei Liang +2

    cs.CVcs.AIarXiv:2103.03454v12021
  49. Active Self-Paced Learning for Cost-Effective and Progressive Face Identification

    Liang Lin, Keze Wang, Deyu Meng +2

    cs.CVarXiv:1701.03555v22017
  50. QuestSim: Human Motion Tracking from Sparse Sensors with Simulated Avatars

    Alexander Winkler, Jungdam Won, Yuting Ye

    cs.CVcs.GRarXiv:2209.09391v12022
  51. Ambient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supervision

    Logesh Kumar Umapathi

    cs.CVcs.AIarXiv:2609.07099v12026
  52. CoSA: Correlation-Guided Change A ttention with Learnable Residual Gating for Remote Sensing Change Detection

    Abdirashid Omar, Jonghyuk Park

    cs.CVarXiv:2609.08914v12026
  53. Single Image to Textured 3D Object Generation in Frequency Domain: From Theory to Pipeline

    Qisen Wang, Yifan Zhao, Jia Li

    cs.CVarXiv:2609.07085v12026
  54. Effective Use of Synthetic Data for Urban Scene Semantic Segmentation

    Fatemeh Sadat Saleh, Mohammad Sadegh Aliakbarian, Mathieu Salzmann +2

    cs.CVarXiv:1807.06132v12018
  55. A Conditional Point Diffusion-Refinement Paradigm for 3D Point Cloud Completion

    Zhaoyang Lyu, Zhifeng Kong, Xudong Xu +2

    cs.CVcs.AIarXiv:2112.03530v42021
  56. Multi-Person Pose Estimation with Enhanced Channel-wise and Spatial Information

    Kai Su, Dongdong Yu, Zhenqi Xu +2

    cs.CVarXiv:1905.03466v12019
  57. TACo: Token-aware Cascade Contrastive Learning for Video-Text Alignment

    Jianwei Yang, Yonatan Bisk, Jianfeng Gao

    cs.CVcs.AIcs.LGarXiv:2108.09980v12021
  58. Real-Time Joint Semantic Segmentation and Depth Estimation Using Asymmetric Annotations

    Vladimir Nekrasov, Thanuja Dharmasiri, Andrew Spek +3

    cs.CVarXiv:1809.04766v22018
  59. Scaling Data Generation in Vision-and-Language Navigation

    Zun Wang, Jialu Li, Yicong Hong +6

    cs.CVcs.AIcs.CLarXiv:2307.15644v22023
  60. Fixed-Rank Representation for Unsupervised Visual Learning

    Risheng Liu, Zhouchen Lin, Fernando De la Torre +1

    cs.CVmath.NAarXiv:1203.2210v22012