Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

2,101 to 2,160 of 18,795

  1. AgenticGen: Reward-Guided Agentic Video Generation for Advertising

    Xingyuan Bu, Chengru Song, Hao Zhou +9

    cs.CVcs.AIcs.CLarXiv:2609.09187v12026
  2. RelGAN: Multi-Domain Image-to-Image Translation via Relative Attributes

    Po-Wei Wu, Yu-Jing Lin, Che-Han Chang +2

    cs.CVarXiv:1908.07269v12019
  3. SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models

    Muyang Li, Yujun Lin, Zhekai Zhang +7

    cs.CVcs.LGarXiv:2411.05007v42024
  4. Explainable Object-induced Action Decision for Autonomous Vehicles

    Yiran Xu, Xiaoyin Yang, Lihang Gong +4

    cs.CVarXiv:2003.09405v12020
  5. SalNet360: Saliency Maps for omni-directional images with CNN

    Rafael Monroy, Sebastian Lutz, Tejo Chalasani +1

    cs.CVarXiv:1709.06505v22017
  6. Adversarial Defense by Restricting the Hidden Space of Deep Neural Networks

    Aamir Mustafa, Salman Khan, Munawar Hayat +3

    cs.CVcs.LGarXiv:1904.00887v42019
  7. Face Attention Network: An Effective Face Detector for the Occluded Faces

    Jianfeng Wang, Ye Yuan, Gang Yu

    cs.CVarXiv:1711.07246v22017
  8. SMD-Nets: Stereo Mixture Density Networks

    Fabio Tosi, Yiyi Liao, Carolin Schmitt +1

    cs.CVarXiv:2104.03866v12021
  9. Representational Continuity for Unsupervised Continual Learning

    Divyam Madaan, Jaehong Yoon, Yuanchun Li +2

    cs.LGcs.CVarXiv:2110.06976v32021
  10. MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis

    Dewei Zhou, You Li, Fan Ma +2

    cs.CVarXiv:2402.05408v22024
  11. Hyperbolic Image-Text Representations

    Karan Desai, Maximilian Nickel, Tanmay Rajpurohit +2

    cs.CVcs.LGarXiv:2304.09172v32023
  12. Quadruplet Network with One-Shot Learning for Fast Visual Object Tracking

    Xingping Dong, Jianbing Shen, Dongming Wu +3

    cs.CVarXiv:1705.07222v32017
  13. Source-Free Domain Adaptive Fundus Image Segmentation with Denoised Pseudo-Labeling

    Cheng Chen, Quande Liu, Yueming Jin +2

    eess.IVcs.CVarXiv:2109.09735v12021
  14. Pores for thought: The use of generative adversarial networks for the stochastic reconstruction of 3D multi-phase electrode microstructures with periodic boundaries

    Andrea Gayon-Lombardo, Lukas Mosser, Nigel P. Brandon +1

    cs.NEcs.CVarXiv:2003.11632v22020
  15. Neural Lumigraph Rendering

    Petr Kellnhofer, Lars Jebe, Andrew Jones +3

    cs.CVcs.GRarXiv:2103.11571v12021
  16. Make-Your-Video: Customized Video Generation Using Textual and Structural Guidance

    Jinbo Xing, Menghan Xia, Yuxin Liu +9

    cs.CVarXiv:2306.00943v12023
  17. One Thing One Click: A Self-Training Approach for Weakly Supervised 3D Semantic Segmentation

    Zhengzhe Liu, Xiaojuan Qi, Chi-Wing Fu

    cs.CVarXiv:2104.02246v42021
  18. Symmetry and Group in Attribute-Object Compositions

    Yong-Lu Li, Yue Xu, Xiaohan Mao +1

    cs.CVcs.LGarXiv:2004.00587v12020
  19. Self-Supervised Monocular 3D Face Reconstruction by Occlusion-Aware Multi-view Geometry Consistency

    Jiaxiang Shang, Tianwei Shen, Shiwei Li +4

    cs.CVarXiv:2007.12494v12020
  20. Unsupervised Multi-source Domain Adaptation Without Access to Source Data

    Sk Miraj Ahmed, Dripta S. Raychaudhuri, Sujoy Paul +2

    cs.LGcs.CVarXiv:2104.01845v12021
  21. Good News, Everyone! Context driven entity-aware captioning for news images

    Ali Furkan Biten, Lluis Gomez, Marçal Rusiñol +1

    cs.CVarXiv:1904.01475v12019
  22. CNN-PS: CNN-based Photometric Stereo for General Non-Convex Surfaces

    Satoshi Ikehata

    cs.CVarXiv:1808.10093v12018
  23. Pedestrian Attribute Recognition: A Survey

    Xiao Wang, Shaofei Zheng, Rui Yang +4

    cs.CVcs.AIcs.LGarXiv:1901.07474v22019
  24. Spatial Fusion GAN for Image Synthesis

    Fangneng Zhan, Hongyuan Zhu, Shijian Lu

    cs.CVarXiv:1812.05840v32018
  25. Optimizing colormaps with consideration for color vision deficiency to enable accurate interpretation of scientific data

    Jamie R. Nuñez, Christopher R. Anderton, Ryan S. Renslow

    cs.CVq-bio.OTarXiv:1712.01662v32017
  26. From Coordinates to Candidate Regions: Temporal Change Localization via Region Selection in Remote Sensing Multimodal LLMs

    Juwan Chung, Sungjune Park, Yeongyun Kim +1

    cs.CVcs.CLarXiv:2609.08391v12026
  27. A comprehensive review of Binary Neural Network

    Chunyu Yuan, Sos S. Agaian

    cs.NEcs.AIcs.CVarXiv:2110.06804v42021
  28. Deep Embedding Convolutional Neural Network for Synthesizing CT Image from T1-Weighted MR Image

    Lei Xiang, Qian Wang, Xiyao Jin +3

    cs.CVarXiv:1709.02073v22017
  29. RenderNet: A deep convolutional network for differentiable rendering from 3D shapes

    Thu Nguyen-Phuoc, Chuan Li, Stephen Balaban +1

    cs.CVarXiv:1806.06575v32018
  30. Re-distributing Biased Pseudo Labels for Semi-supervised Semantic Segmentation: A Baseline Investigation

    Ruifei He, Jihan Yang, Xiaojuan Qi

    cs.CVarXiv:2107.11279v22021
  31. Novel Methods for Catheter and Guidewire Segmentation in X-ray Fluoroscopy under a Federated Learning Setting

    Chayun Kongtongvattana

    cs.CVcs.AIarXiv:2609.06876v12026
  32. Probabilistic 3D Multi-Modal, Multi-Object Tracking for Autonomous Driving

    Hsu-kuang Chiu, Jie Li, Rares Ambrus +1

    cs.CVcs.ROarXiv:2012.13755v22020
  33. Engaging Image Captioning Via Personality

    Kurt Shuster, Samuel Humeau, Hexiang Hu +2

    cs.CVcs.AIcs.CLarXiv:1810.10665v22018
  34. Spatial Attention Deep Net with Partial PSO for Hierarchical Hybrid Hand Pose Estimation

    Qi Ye, Shanxin Yuan, Tae-Kyun Kim

    cs.CVarXiv:1604.03334v22016
  35. TextScanner: Reading Characters in Order for Robust Scene Text Recognition

    Zhaoyi Wan, Minghang He, Haoran Chen +2

    cs.CVcs.CLcs.LGarXiv:1912.12422v22019
  36. WarpNet: Weakly Supervised Matching for Single-view Reconstruction

    Angjoo Kanazawa, David W. Jacobs, Manmohan Chandraker

    cs.CVarXiv:1604.05592v22016
  37. 200x Low-dose PET Reconstruction using Deep Learning

    Junshen Xu, Enhao Gong, John Pauly +1

    cs.CVarXiv:1712.04119v12017
  38. Pose Guided Human Video Generation

    Ceyuan Yang, Zhe Wang, Xinge Zhu +3

    cs.CVarXiv:1807.11152v12018
  39. Learning to Infer and Execute 3D Shape Programs

    Yonglong Tian, Andrew Luo, Xingyuan Sun +4

    cs.CVcs.AIcs.GRarXiv:1901.02875v32019
  40. Pavement Image Datasets: A New Benchmark Dataset to Classify and Densify Pavement Distresses

    Hamed Majidifard, Peng Jin, Yaw Adu-Gyamfi +1

    cs.CVcs.LGstat.MLarXiv:1910.11123v22019
  41. Emo-DVS: A Multimodal Benchmark for Privacy-Aware Emotion Recognition with Event Cameras

    Jiaqi Chen, Qinfu Xu, Hao Zhuang +1

    cs.CVcs.AIarXiv:2609.06928v12026
  42. HED-UNet: Combined Segmentation and Edge Detection for Monitoring the Antarctic Coastline

    Konrad Heidler, Lichao Mou, Celia Baumhoer +2

    cs.CVeess.IVarXiv:2103.01849v12021
  43. Deep Fitting Degree Scoring Network for Monocular 3D Object Detection

    Lijie Liu, Jiwen Lu, Chunjing Xu +2

    cs.CVarXiv:1904.12681v22019
  44. Multi-Objective Interpolation Training for Robustness to Label Noise

    Diego Ortego, Eric Arazo, Paul Albert +2

    cs.CVarXiv:2012.04462v22020
  45. HuMMan: Multi-Modal 4D Human Dataset for Versatile Sensing and Modeling

    Zhongang Cai, Daxuan Ren, Ailing Zeng +12

    cs.CVarXiv:2204.13686v22022
  46. iVideoGPT: Interactive VideoGPTs are Scalable World Models

    Jialong Wu, Shaofeng Yin, Ningya Feng +4

    cs.CVcs.LGcs.ROarXiv:2405.15223v32024
  47. Less is More: Lighter and Faster Deep Neural Architecture for Tomato Leaf Disease Classification

    Sabbir Ahmed, Md. Bakhtiar Hasan, Tasnim Ahmed +2

    cs.CVcs.LGarXiv:2109.02394v22021
  48. Explicit Attention-Enhanced Fusion for RGB-Thermal Perception Tasks

    Mingjian Liang, Junjie Hu, Chenyu Bao +3

    cs.CVarXiv:2303.15710v12023
  49. Knowledge Distillation with Adversarial Samples Supporting Decision Boundary

    Byeongho Heo, Minsik Lee, Sangdoo Yun +1

    cs.LGcs.CVstat.MLarXiv:1805.05532v42018
  50. Hierarchical interpretations for neural network predictions

    Chandan Singh, W. James Murdoch, Bin Yu

    cs.LGcs.AIcs.CLarXiv:1806.05337v22018
  51. End-to-end Multi-Modal Multi-Task Vehicle Control for Self-Driving Cars with Visual Perception

    Zhengyuan Yang, Yixuan Zhang, Jerry Yu +2

    cs.CVarXiv:1801.06734v22018
  52. Learnable Fourier Features for Multi-Dimensional Spatial Positional Encoding

    Yang Li, Si Si, Gang Li +2

    cs.LGcs.AIcs.CVarXiv:2106.02795v32021
  53. Compact Transformer Tracker with Correlative Masked Modeling

    Zikai Song, Run Luo, Junqing Yu +2

    cs.CVarXiv:2301.10938v12023
  54. A DeNoising FPN With Transformer R-CNN for Tiny Object Detection

    Hou-I Liu, Yu-Wen Tseng, Kai-Cheng Chang +3

    cs.CVarXiv:2406.05755v42024
  55. Evolution of Multimodal Question Answering: From Modality-Adaptive Extraction to Unified Language Representation

    Abdullah Al Shafi

    cs.CLcs.CVarXiv:2609.08896v12026
  56. Deep Imitative Models for Flexible Inference, Planning, and Control

    Nicholas Rhinehart, Rowan McAllister, Sergey Levine

    cs.LGcs.AIcs.CVarXiv:1810.06544v42018
  57. Uni-Perceiver: Pre-training Unified Architecture for Generic Perception for Zero-shot and Few-shot Tasks

    Xizhou Zhu, Jinguo Zhu, Hao Li +5

    cs.CVarXiv:2112.01522v12021
  58. Dynamic Coarse-to-Fine Learning for Oriented Tiny Object Detection

    Chang Xu, Jian Ding, Jinwang Wang +4

    cs.CVarXiv:2304.08876v12023
  59. Visualizing Deep Networks by Optimizing with Integrated Gradients

    Zhongang Qi, Saeed Khorram, Li Fuxin

    cs.CVarXiv:1905.00954v22019
  60. Asymmetric Co-Teaching for Unsupervised Cross Domain Person Re-Identification

    Fengxiang Yang, Ke Li, Zhun Zhong +7

    cs.CVarXiv:1912.01349v12019