Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
2,101 to 2,160 of 18,795
AgenticGen: Reward-Guided Agentic Video Generation for Advertising
Xingyuan Bu, Chengru Song, Hao Zhou +9
cs.CVcs.AIcs.CLarXiv:2609.09187v12026RelGAN: Multi-Domain Image-to-Image Translation via Relative Attributes
Po-Wei Wu, Yu-Jing Lin, Che-Han Chang +2
cs.CVarXiv:1908.07269v12019SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models
Muyang Li, Yujun Lin, Zhekai Zhang +7
cs.CVcs.LGarXiv:2411.05007v42024Explainable Object-induced Action Decision for Autonomous Vehicles
Yiran Xu, Xiaoyin Yang, Lihang Gong +4
cs.CVarXiv:2003.09405v12020SalNet360: Saliency Maps for omni-directional images with CNN
Rafael Monroy, Sebastian Lutz, Tejo Chalasani +1
cs.CVarXiv:1709.06505v22017Adversarial Defense by Restricting the Hidden Space of Deep Neural Networks
Aamir Mustafa, Salman Khan, Munawar Hayat +3
cs.CVcs.LGarXiv:1904.00887v42019Face Attention Network: An Effective Face Detector for the Occluded Faces
Jianfeng Wang, Ye Yuan, Gang Yu
cs.CVarXiv:1711.07246v22017SMD-Nets: Stereo Mixture Density Networks
Fabio Tosi, Yiyi Liao, Carolin Schmitt +1
cs.CVarXiv:2104.03866v12021Representational Continuity for Unsupervised Continual Learning
Divyam Madaan, Jaehong Yoon, Yuanchun Li +2
cs.LGcs.CVarXiv:2110.06976v32021MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis
Dewei Zhou, You Li, Fan Ma +2
cs.CVarXiv:2402.05408v22024Hyperbolic Image-Text Representations
Karan Desai, Maximilian Nickel, Tanmay Rajpurohit +2
cs.CVcs.LGarXiv:2304.09172v32023Quadruplet Network with One-Shot Learning for Fast Visual Object Tracking
Xingping Dong, Jianbing Shen, Dongming Wu +3
cs.CVarXiv:1705.07222v32017Source-Free Domain Adaptive Fundus Image Segmentation with Denoised Pseudo-Labeling
Cheng Chen, Quande Liu, Yueming Jin +2
eess.IVcs.CVarXiv:2109.09735v12021Pores for thought: The use of generative adversarial networks for the stochastic reconstruction of 3D multi-phase electrode microstructures with periodic boundaries
Andrea Gayon-Lombardo, Lukas Mosser, Nigel P. Brandon +1
cs.NEcs.CVarXiv:2003.11632v22020Neural Lumigraph Rendering
Petr Kellnhofer, Lars Jebe, Andrew Jones +3
cs.CVcs.GRarXiv:2103.11571v12021Make-Your-Video: Customized Video Generation Using Textual and Structural Guidance
Jinbo Xing, Menghan Xia, Yuxin Liu +9
cs.CVarXiv:2306.00943v12023One Thing One Click: A Self-Training Approach for Weakly Supervised 3D Semantic Segmentation
Zhengzhe Liu, Xiaojuan Qi, Chi-Wing Fu
cs.CVarXiv:2104.02246v42021Symmetry and Group in Attribute-Object Compositions
Yong-Lu Li, Yue Xu, Xiaohan Mao +1
cs.CVcs.LGarXiv:2004.00587v12020Self-Supervised Monocular 3D Face Reconstruction by Occlusion-Aware Multi-view Geometry Consistency
Jiaxiang Shang, Tianwei Shen, Shiwei Li +4
cs.CVarXiv:2007.12494v12020Unsupervised Multi-source Domain Adaptation Without Access to Source Data
Sk Miraj Ahmed, Dripta S. Raychaudhuri, Sujoy Paul +2
cs.LGcs.CVarXiv:2104.01845v12021Good News, Everyone! Context driven entity-aware captioning for news images
Ali Furkan Biten, Lluis Gomez, Marçal Rusiñol +1
cs.CVarXiv:1904.01475v12019CNN-PS: CNN-based Photometric Stereo for General Non-Convex Surfaces
Satoshi Ikehata
cs.CVarXiv:1808.10093v12018Pedestrian Attribute Recognition: A Survey
Xiao Wang, Shaofei Zheng, Rui Yang +4
cs.CVcs.AIcs.LGarXiv:1901.07474v22019Spatial Fusion GAN for Image Synthesis
Fangneng Zhan, Hongyuan Zhu, Shijian Lu
cs.CVarXiv:1812.05840v32018Optimizing colormaps with consideration for color vision deficiency to enable accurate interpretation of scientific data
Jamie R. Nuñez, Christopher R. Anderton, Ryan S. Renslow
cs.CVq-bio.OTarXiv:1712.01662v32017From Coordinates to Candidate Regions: Temporal Change Localization via Region Selection in Remote Sensing Multimodal LLMs
Juwan Chung, Sungjune Park, Yeongyun Kim +1
cs.CVcs.CLarXiv:2609.08391v12026A comprehensive review of Binary Neural Network
Chunyu Yuan, Sos S. Agaian
cs.NEcs.AIcs.CVarXiv:2110.06804v42021Deep Embedding Convolutional Neural Network for Synthesizing CT Image from T1-Weighted MR Image
Lei Xiang, Qian Wang, Xiyao Jin +3
cs.CVarXiv:1709.02073v22017RenderNet: A deep convolutional network for differentiable rendering from 3D shapes
Thu Nguyen-Phuoc, Chuan Li, Stephen Balaban +1
cs.CVarXiv:1806.06575v32018Re-distributing Biased Pseudo Labels for Semi-supervised Semantic Segmentation: A Baseline Investigation
Ruifei He, Jihan Yang, Xiaojuan Qi
cs.CVarXiv:2107.11279v22021Novel Methods for Catheter and Guidewire Segmentation in X-ray Fluoroscopy under a Federated Learning Setting
Chayun Kongtongvattana
cs.CVcs.AIarXiv:2609.06876v12026Probabilistic 3D Multi-Modal, Multi-Object Tracking for Autonomous Driving
Hsu-kuang Chiu, Jie Li, Rares Ambrus +1
cs.CVcs.ROarXiv:2012.13755v22020Engaging Image Captioning Via Personality
Kurt Shuster, Samuel Humeau, Hexiang Hu +2
cs.CVcs.AIcs.CLarXiv:1810.10665v22018Spatial Attention Deep Net with Partial PSO for Hierarchical Hybrid Hand Pose Estimation
Qi Ye, Shanxin Yuan, Tae-Kyun Kim
cs.CVarXiv:1604.03334v22016TextScanner: Reading Characters in Order for Robust Scene Text Recognition
Zhaoyi Wan, Minghang He, Haoran Chen +2
cs.CVcs.CLcs.LGarXiv:1912.12422v22019WarpNet: Weakly Supervised Matching for Single-view Reconstruction
Angjoo Kanazawa, David W. Jacobs, Manmohan Chandraker
cs.CVarXiv:1604.05592v22016200x Low-dose PET Reconstruction using Deep Learning
Junshen Xu, Enhao Gong, John Pauly +1
cs.CVarXiv:1712.04119v12017Pose Guided Human Video Generation
Ceyuan Yang, Zhe Wang, Xinge Zhu +3
cs.CVarXiv:1807.11152v12018Learning to Infer and Execute 3D Shape Programs
Yonglong Tian, Andrew Luo, Xingyuan Sun +4
cs.CVcs.AIcs.GRarXiv:1901.02875v32019Pavement Image Datasets: A New Benchmark Dataset to Classify and Densify Pavement Distresses
Hamed Majidifard, Peng Jin, Yaw Adu-Gyamfi +1
cs.CVcs.LGstat.MLarXiv:1910.11123v22019Emo-DVS: A Multimodal Benchmark for Privacy-Aware Emotion Recognition with Event Cameras
Jiaqi Chen, Qinfu Xu, Hao Zhuang +1
cs.CVcs.AIarXiv:2609.06928v12026HED-UNet: Combined Segmentation and Edge Detection for Monitoring the Antarctic Coastline
Konrad Heidler, Lichao Mou, Celia Baumhoer +2
cs.CVeess.IVarXiv:2103.01849v12021Deep Fitting Degree Scoring Network for Monocular 3D Object Detection
Lijie Liu, Jiwen Lu, Chunjing Xu +2
cs.CVarXiv:1904.12681v22019Multi-Objective Interpolation Training for Robustness to Label Noise
Diego Ortego, Eric Arazo, Paul Albert +2
cs.CVarXiv:2012.04462v22020HuMMan: Multi-Modal 4D Human Dataset for Versatile Sensing and Modeling
Zhongang Cai, Daxuan Ren, Ailing Zeng +12
cs.CVarXiv:2204.13686v22022iVideoGPT: Interactive VideoGPTs are Scalable World Models
Jialong Wu, Shaofeng Yin, Ningya Feng +4
cs.CVcs.LGcs.ROarXiv:2405.15223v32024Less is More: Lighter and Faster Deep Neural Architecture for Tomato Leaf Disease Classification
Sabbir Ahmed, Md. Bakhtiar Hasan, Tasnim Ahmed +2
cs.CVcs.LGarXiv:2109.02394v22021Explicit Attention-Enhanced Fusion for RGB-Thermal Perception Tasks
Mingjian Liang, Junjie Hu, Chenyu Bao +3
cs.CVarXiv:2303.15710v12023Knowledge Distillation with Adversarial Samples Supporting Decision Boundary
Byeongho Heo, Minsik Lee, Sangdoo Yun +1
cs.LGcs.CVstat.MLarXiv:1805.05532v42018Hierarchical interpretations for neural network predictions
Chandan Singh, W. James Murdoch, Bin Yu
cs.LGcs.AIcs.CLarXiv:1806.05337v22018End-to-end Multi-Modal Multi-Task Vehicle Control for Self-Driving Cars with Visual Perception
Zhengyuan Yang, Yixuan Zhang, Jerry Yu +2
cs.CVarXiv:1801.06734v22018Learnable Fourier Features for Multi-Dimensional Spatial Positional Encoding
Yang Li, Si Si, Gang Li +2
cs.LGcs.AIcs.CVarXiv:2106.02795v32021Compact Transformer Tracker with Correlative Masked Modeling
Zikai Song, Run Luo, Junqing Yu +2
cs.CVarXiv:2301.10938v12023A DeNoising FPN With Transformer R-CNN for Tiny Object Detection
Hou-I Liu, Yu-Wen Tseng, Kai-Cheng Chang +3
cs.CVarXiv:2406.05755v42024Evolution of Multimodal Question Answering: From Modality-Adaptive Extraction to Unified Language Representation
Abdullah Al Shafi
cs.CLcs.CVarXiv:2609.08896v12026Deep Imitative Models for Flexible Inference, Planning, and Control
Nicholas Rhinehart, Rowan McAllister, Sergey Levine
cs.LGcs.AIcs.CVarXiv:1810.06544v42018Uni-Perceiver: Pre-training Unified Architecture for Generic Perception for Zero-shot and Few-shot Tasks
Xizhou Zhu, Jinguo Zhu, Hao Li +5
cs.CVarXiv:2112.01522v12021Dynamic Coarse-to-Fine Learning for Oriented Tiny Object Detection
Chang Xu, Jian Ding, Jinwang Wang +4
cs.CVarXiv:2304.08876v12023Visualizing Deep Networks by Optimizing with Integrated Gradients
Zhongang Qi, Saeed Khorram, Li Fuxin
cs.CVarXiv:1905.00954v22019Asymmetric Co-Teaching for Unsupervised Cross Domain Person Re-Identification
Fengxiang Yang, Ke Li, Zhun Zhong +7
cs.CVarXiv:1912.01349v12019