Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

8,941 to 9,000 of 18,866

  1. Image Generation from Layout

    Bo Zhao, Lili Meng, Weidong Yin +1

    cs.CVeess.IVarXiv:1811.11389v32018
  2. MSeg: A Composite Dataset for Multi-domain Semantic Segmentation

    John Lambert, Zhuang Liu, Ozan Sener +2

    cs.CVarXiv:2112.13762v12021
  3. UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild

    Can Qin, Shu Zhang, Ning Yu +10

    cs.CVcs.AIarXiv:2305.11147v32023
  4. A Transformer-Based Feature Segmentation and Region Alignment Method For UAV-View Geo-Localization

    Ming Dai, Jianhong Hu, Jiedong Zhuang +1

    cs.CVcs.AIarXiv:2201.09206v12022
  5. Video Propagation Networks

    Varun Jampani, Raghudeep Gadde, Peter V. Gehler

    cs.CVarXiv:1612.05478v32016
  6. Stabilizing Differentiable Architecture Search via Perturbation-based Regularization

    Xiangning Chen, Cho-Jui Hsieh

    cs.LGcs.CVstat.MLarXiv:2002.05283v32020
  7. Implicit Semantic Data Augmentation for Deep Networks

    Yulin Wang, Xuran Pan, Shiji Song +3

    cs.CVcs.LGstat.MLarXiv:1909.12220v52019
  8. Superpixel Segmentation with Fully Convolutional Networks

    Fengting Yang, Qian Sun, Hailin Jin +1

    cs.CVarXiv:2003.12929v12020
  9. MedKLIP: Medical Knowledge Enhanced Language-Image Pre-Training in Radiology

    Chaoyi Wu, Xiaoman Zhang, Ya Zhang +2

    eess.IVcs.CLcs.CVarXiv:2301.02228v32023
  10. EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion Model

    Xinya Ji, Hang Zhou, Kaisiyuan Wang +4

    cs.CVarXiv:2205.15278v32022
  11. To Generate or Not? Safety-Driven Unlearned Diffusion Models Are Still Easy To Generate Unsafe Images ... For Now

    Yimeng Zhang, Jinghan Jia, Xin Chen +5

    cs.CVarXiv:2310.11868v42023
  12. End-to-end Concept Word Detection for Video Captioning, Retrieval, and Question Answering

    Youngjae Yu, Hyungjin Ko, Jongwook Choi +1

    cs.CVarXiv:1610.02947v32016
  13. Deformable Shape Completion with Graph Convolutional Autoencoders

    Or Litany, Alex Bronstein, Michael Bronstein +1

    cs.CVarXiv:1712.00268v42017
  14. ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models

    Shaghayegh Kolli, Sina Emami, Moreno D'Incà +4

    cs.CVcs.CLarXiv:2608.29847v12026
  15. Spatially-sparse convolutional neural networks

    Benjamin Graham

    cs.CVcs.NEarXiv:1409.6070v12014
  16. Deep Representation Learning on Long-tailed Data: A Learnable Embedding Augmentation Perspective

    Jialun Liu, Yifan Sun, Chuchu Han +2

    cs.CVarXiv:2002.10826v32020
  17. Spatiotemporal Pyramid Network for Video Action Recognition

    Yunbo Wang, Mingsheng Long, Jianmin Wang +1

    cs.CVarXiv:1903.01038v12019
  18. StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text

    Roberto Henschel, Levon Khachatryan, Hayk Poghosyan +5

    cs.CVcs.AIcs.CLarXiv:2403.14773v22024
  19. Few-Shot Segmentation Without Meta-Learning: A Good Transductive Inference Is All You Need?

    Malik Boudiaf, Hoel Kervadec, Ziko Imtiaz Masud +3

    cs.CVarXiv:2012.06166v22020
  20. CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs

    Zirui Wang, Mengzhou Xia, Luxi He +10

    cs.CLcs.CVarXiv:2406.18521v12024
  21. Natural Synthetic Anomalies for Self-Supervised Anomaly Detection and Localization

    Hannah M. Schlüter, Jeremy Tan, Benjamin Hou +1

    cs.CVarXiv:2109.15222v32021
  22. Diffusion Hyperfeatures: Searching Through Time and Space for Semantic Correspondence

    Grace Luo, Lisa Dunlap, Dong Huk Park +2

    cs.CVarXiv:2305.14334v22023
  23. Targeting Ultimate Accuracy: Face Recognition via Deep Embedding

    Jingtuo Liu, Yafeng Deng, Tao Bai +2

    cs.CVarXiv:1506.07310v42015
  24. Progressive Domain Expansion Network for Single Domain Generalization

    Lei Li, Ke Gao, Juan Cao +6

    cs.CVarXiv:2103.16050v12021
  25. Tracking Everything Everywhere All at Once

    Qianqian Wang, Yen-Yu Chang, Ruojin Cai +4

    cs.CVarXiv:2306.05422v22023
  26. SoccerNet-v2: A Dataset and Benchmarks for Holistic Understanding of Broadcast Soccer Videos

    Adrien Deliège, Anthony Cioppa, Silvio Giancola +6

    cs.CVarXiv:2011.13367v32020
  27. 3D-LaneNet: End-to-End 3D Multiple Lane Detection

    Noa Garnett, Rafi Cohen, Tomer Pe'er +2

    cs.CVcs.LGcs.ROarXiv:1811.10203v32018
  28. ShowUI: One Vision-Language-Action Model for GUI Visual Agent

    Kevin Qinghong Lin, Linjie Li, Difei Gao +6

    cs.CVcs.AIcs.CLarXiv:2411.17465v12024
  29. DDP: Diffusion Model for Dense Visual Prediction

    Yuanfeng Ji, Zhe Chen, Enze Xie +6

    cs.CVarXiv:2303.17559v22023
  30. Cross-domain Object Detection through Coarse-to-Fine Feature Adaptation

    Yangtao Zheng, Di Huang, Songtao Liu +1

    cs.CVarXiv:2003.10275v12020
  31. Logit Standardization in Knowledge Distillation

    Shangquan Sun, Wenqi Ren, Jingzhi Li +2

    cs.CVarXiv:2403.01427v12024
  32. Co-Fusion: Real-time Segmentation, Tracking and Fusion of Multiple Objects

    Martin Rünz, Lourdes Agapito

    cs.CVarXiv:1706.06629v12017
  33. R2GenGPT: Radiology Report Generation with Frozen LLMs

    Zhanyu Wang, Lingqiao Liu, Lei Wang +1

    cs.CVarXiv:2309.09812v22023
  34. DDD17: End-To-End DAVIS Driving Dataset

    Jonathan Binas, Daniel Neil, Shih-Chii Liu +1

    cs.CVarXiv:1711.01458v12017
  35. SparseFool: a few pixels make a big difference

    Apostolos Modas, Seyed-Mohsen Moosavi-Dezfooli, Pascal Frossard

    cs.CVcs.CRcs.LGarXiv:1811.02248v42018
  36. Integrating Triaxial IMU Sensors and Ensemble Learning for Effective Parkinson Disease Severity Classification

    Rehan Khan, Muhammad Junaid Asif, Rana Fayyaz Ahmad

    cs.AIcs.CVarXiv:2608.28602v12026
  37. Language Embedded 3D Gaussians for Open-Vocabulary Scene Understanding

    Jin-Chuan Shi, Miao Wang, Hao-Bin Duan +1

    cs.CVcs.GRarXiv:2311.18482v12023
  38. MedTVL: Harnessing Vision and Language for Medical Time Series Classification

    Jiexia Ye, Jia Li, Fugee Tsung

    cs.AIcs.CVarXiv:2608.28605v12026
  39. Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model

    Feng Liu, Shiwei Zhang, Xiaofeng Wang +6

    cs.CVarXiv:2411.19108v22024
  40. Local Learning Matters: Rethinking Data Heterogeneity in Federated Learning

    Matias Mendieta, Taojiannan Yang, Pu Wang +3

    cs.LGcs.CVcs.DCarXiv:2111.14213v32021
  41. Depth-based hand pose estimation: methods, data, and challenges

    James Steven Supancic, Gregory Rogez, Yi Yang +2

    cs.CVarXiv:1504.06378v22015
  42. The Potential of Haptic Foundation Models

    Jianquan Wang, Haiwei Dong, Abdulmotaleb El Saddik

    cs.ROcs.CVcs.MMarXiv:2608.28664v12026
  43. Generating Holistic 3D Human Motion from Speech

    Hongwei Yi, Hualin Liang, Yifei Liu +5

    cs.CVcs.GRarXiv:2212.04420v22022
  44. Multi-exposure HDR Imaging: A Review of Pixel-level and Feature-level Reconstruction Methods

    Qian Tao, Wei Wang, Chaobing Zheng +1

    cs.CVarXiv:2608.28674v12026
  45. Query-Dependent Video Representation for Moment Retrieval and Highlight Detection

    WonJun Moon, Sangeek Hyun, SangUk Park +2

    cs.CVcs.AIarXiv:2303.13874v12023
  46. Deep Reinforcement Learning in Computer Vision: A Comprehensive Survey

    Ngan Le, Vidhiwar Singh Rathour, Kashu Yamazaki +2

    cs.CVcs.AIarXiv:2108.11510v12021
  47. Review of Large Vision Models and Visual Prompt Engineering

    Jiaqi Wang, Zhengliang Liu, Lin Zhao +18

    cs.CVcs.AIarXiv:2307.00855v12023
  48. A Bayesian Data Augmentation Approach for Learning Deep Models

    Toan Tran, Trung Pham, Gustavo Carneiro +2

    cs.CVcs.LGarXiv:1710.10564v12017
  49. Driving Policy Transfer via Modularity and Abstraction

    Matthias Müller, Alexey Dosovitskiy, Bernard Ghanem +1

    cs.ROcs.CVcs.LGarXiv:1804.09364v32018
  50. Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning

    Yingdong Hu, Fanqi Lin, Tong Zhang +2

    cs.ROcs.AIcs.CLarXiv:2311.17842v22023
  51. M$^2$BEV: Multi-Camera Joint 3D Detection and Segmentation with Unified Birds-Eye View Representation

    Enze Xie, Zhiding Yu, Daquan Zhou +5

    cs.CVarXiv:2204.05088v22022
  52. PI-RCNN: An Efficient Multi-sensor 3D Object Detector with Point-based Attentive Cont-conv Fusion Module

    Liang Xie, Chao Xiang, Zhengxu Yu +4

    cs.CVarXiv:1911.06084v32019
  53. Referring Image Segmentation via Cross-Modal Progressive Comprehension

    Shaofei Huang, Tianrui Hui, Si Liu +5

    cs.CVcs.CLarXiv:2010.00514v12020
  54. BundleSDF: Neural 6-DoF Tracking and 3D Reconstruction of Unknown Objects

    Bowen Wen, Jonathan Tremblay, Valts Blukis +6

    cs.CVcs.AIcs.GRarXiv:2303.14158v12023
  55. When Image Denoising Meets High-Level Vision Tasks: A Deep Learning Approach

    Ding Liu, Bihan Wen, Xianming Liu +2

    cs.CVarXiv:1706.04284v32017
  56. Joint Monocular 3D Vehicle Detection and Tracking

    Hou-Ning Hu, Qi-Zhi Cai, Dequan Wang +5

    cs.CVarXiv:1811.10742v32018
  57. Fooling Neural Network Interpretations via Adversarial Model Manipulation

    Juyeon Heo, Sunghwan Joo, Taesup Moon

    cs.LGcs.AIcs.CVarXiv:1902.02041v32019
  58. Actor-Centric Relation Network

    Chen Sun, Abhinav Shrivastava, Carl Vondrick +3

    cs.CVarXiv:1807.10982v12018
  59. Fracture Detection in Pediatric Wrist Trauma X-ray Images Using YOLOv8 Algorithm

    Rui-Yang Ju, Weiming Cai

    cs.CVarXiv:2304.05071v52023
  60. Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding

    Xintong Wang, Jingheng Pan, Liang Ding +1

    cs.CVcs.AIcs.CLarXiv:2403.18715v22024