Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
9,601 to 9,660 of 18,776
Illuminating Pedestrians via Simultaneous Detection & Segmentation
Garrick Brazil, Xi Yin, Xiaoming Liu
cs.CVarXiv:1706.08564v12017Visual Question Answering: Datasets, Algorithms, and Future Challenges
Kushal Kafle, Christopher Kanan
cs.CVcs.AIcs.CLarXiv:1610.01465v42016ET-Net: A Generic Edge-aTtention Guidance Network for Medical Image Segmentation
Zhijie Zhang, Huazhu Fu, Hang Dai +3
cs.CVarXiv:1907.10936v12019Summaries:한국어Adversarial Diversity and Hard Positive Generation
Andras Rozsa, Ethan M. Rudd, Terrance E. Boult
cs.CVarXiv:1605.01775v22016TransNet V2: An effective deep network architecture for fast shot transition detection
Tomáš Souček, Jakub Lokoč
cs.CVarXiv:2008.04838v12020Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Jianwei Yang, Hao Zhang, Feng Li +3
cs.CVcs.AIcs.CLarXiv:2310.11441v22023MemSeg: A semi-supervised method for image surface defect detection using differences and commonalities
Minghui Yang, Peng Wu, Jing Liu +1
cs.CVarXiv:2205.00908v12022X2CT-GAN: Reconstructing CT from Biplanar X-Rays with Generative Adversarial Networks
Xingde Ying, Heng Guo, Kai Ma +3
eess.IVcs.CVarXiv:1905.06902v12019Structured Feature Learning for Pose Estimation
Xiao Chu, Wanli Ouyang, Hongsheng Li +1
cs.CVarXiv:1603.09065v12016Affect Analysis in-the-wild: Valence-Arousal, Expressions, Action Units and a Unified Framework
Dimitrios Kollias, Stefanos Zafeiriou
cs.CVcs.AIcs.LGarXiv:2103.15792v12021Gotta Go Fast When Generating Data with Score-Based Models
Alexia Jolicoeur-Martineau, Ke Li, Rémi Piché-Taillefer +2
cs.LGcs.CVmath.OCarXiv:2105.14080v12021Self-Supervised Video Hashing with Hierarchical Binary Auto-encoder
Jingkuan Song, Hanwang Zhang, Xiangpeng Li +3
cs.CVarXiv:1802.02305v12018Do Datasets Have Politics? Disciplinary Values in Computer Vision Dataset Development
Morgan Klaus Scheuerman, Emily Denton, Alex Hanna
cs.CVcs.HCarXiv:2108.04308v22021A Hybrid Deep Learning Architecture for Privacy-Preserving Mobile Analytics
Seyed Ali Osia, Ali Shahin Shamsabadi, Sina Sajadmanesh +5
cs.LGcs.CVarXiv:1703.02952v72017Kernel Methods on Riemannian Manifolds with Gaussian RBF Kernels
Sadeep Jayasumana, Richard Hartley, Mathieu Salzmann +2
cs.CVarXiv:1412.0265v22014Learning by Aligning: Visible-Infrared Person Re-identification using Cross-Modal Correspondences
Hyunjong Park, Sanghoon Lee, Junghyup Lee +1
cs.CVarXiv:2108.07422v12021FixBi: Bridging Domain Spaces for Unsupervised Domain Adaptation
Jaemin Na, Heechul Jung, Hyung Jin Chang +1
cs.CVarXiv:2011.09230v22020Two-Stream Network for Sign Language Recognition and Translation
Yutong Chen, Ronglai Zuo, Fangyun Wei +3
cs.CVarXiv:2211.01367v22022See Better Before Looking Closer: Weakly Supervised Data Augmentation Network for Fine-Grained Visual Classification
Tao Hu, Honggang Qi, Qingming Huang +1
cs.CVarXiv:1901.09891v22019FourLLIE: Boosting Low-Light Image Enhancement by Fourier Frequency Information
Chenxi Wang, Hongjun Wu, Zhi Jin
cs.CVeess.IVarXiv:2308.03033v12023Rethinking on Multi-Stage Networks for Human Pose Estimation
Wenbo Li, Zhicheng Wang, Binyi Yin +7
cs.CVarXiv:1901.00148v42019RIO: 3D Object Instance Re-Localization in Changing Indoor Environments
Johanna Wald, Armen Avetisyan, Nassir Navab +2
cs.CVarXiv:1908.06109v12019ACRONYM: A Large-Scale Grasp Dataset Based on Simulation
Clemens Eppner, Arsalan Mousavian, Dieter Fox
cs.ROcs.CVarXiv:2011.09584v12020NumBench: Diagnosing Counting Failures in Text-to-Image Models
Sandeep Wadhwa, Mayank Vatsa, Richa Singh +2
cs.CVcs.DBarXiv:2608.28206v12026Crowd counting via scale-adaptive convolutional neural network
Lu Zhang, Miaojing Shi, Qiaobo Chen
cs.CVarXiv:1711.04433v42017The Sound of Motions
Hang Zhao, Chuang Gan, Wei-Chiu Ma +1
cs.CVcs.SDeess.ASarXiv:1904.05979v12019Fast End-to-End Trainable Guided Filter
Huikai Wu, Shuai Zheng, Junge Zhang +1
cs.CVarXiv:1803.05619v22018Universal Instance Perception as Object Discovery and Retrieval
Bin Yan, Yi Jiang, Jiannan Wu +4
cs.CVarXiv:2303.06674v22023Self-Chained Image-Language Model for Video Localization and Question Answering
Shoubin Yu, Jaemin Cho, Prateek Yadav +1
cs.CVcs.AIcs.CLarXiv:2305.06988v220233D Reconstruction with Spatial Memory
Hengyi Wang, Lourdes Agapito
cs.CVarXiv:2408.16061v12024End-to-End Human Object Interaction Detection with HOI Transformer
Cheng Zou, Bohan Wang, Yue Hu +8
cs.CVarXiv:2103.04503v12021Pinwheel-shaped Convolution and Scale-based Dynamic Loss for Infrared Small Target Detection
Jiangnan Yang, Shuangli Liu, Jingjun Wu +3
cs.CVarXiv:2412.16986v12024DensityKV: Density-Guided KV Cache Compression for Long Video Generation
Wenqu Zhao, Xuemin Chi, Xin Zhang +6
cs.CVarXiv:2608.27922v12026Equivariant Multi-Modality Image Fusion
Zixiang Zhao, Haowen Bai, Jiangshe Zhang +6
cs.CVarXiv:2305.11443v22023Towards Visually Explaining Variational Autoencoders
Wenqian Liu, Runze Li, Meng Zheng +5
cs.CVcs.LGarXiv:1911.07389v72019Federated Learning for Medical Applications: A Taxonomy, Current Trends, Challenges, and Future Research Directions
Ashish Rauniyar, Desta Haileselassie Hagos, Debesh Jha +4
cs.LGcs.CRcs.CVarXiv:2208.03392v52022Efficient and Accurate Approximations of Nonlinear Convolutional Networks
Xiangyu Zhang, Jianhua Zou, Xiang Ming +2
cs.CVarXiv:1411.4229v12014HyperDreamBooth: HyperNetworks for Fast Personalization of Text-to-Image Models
Nataniel Ruiz, Yuanzhen Li, Varun Jampani +6
cs.CVcs.AIcs.GRarXiv:2307.06949v22023Geometric Feature-Based Facial Expression Recognition in Image Sequences Using Multi-Class AdaBoost and Support Vector Machines
Deepak Ghimire, Joonwhoan Lee
cs.CVarXiv:1604.03225v12016Climate Physics Dynamic Matching
Gurjeet Sangra Singh, Frantzeska Lavda, Alexandros Kalousis
stat.APcs.CVcs.LOarXiv:2608.26907v12026ICDAR2019 Robust Reading Challenge on Arbitrary-Shaped Text (RRC-ArT)
Chee-Kheng Chng, Yuliang Liu, Yipeng Sun +11
cs.CVarXiv:1909.07145v12019Self-supervised Learning with Geometric Constraints in Monocular Video: Connecting Flow, Depth, and Camera
Yuhua Chen, Cordelia Schmid, Cristian Sminchisescu
cs.CVarXiv:1907.05820v22019MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier +29
cs.CVcs.CLcs.LGarXiv:2403.09611v42024AutoAssign: Differentiable Label Assignment for Dense Object Detection
Benjin Zhu, Jianfeng Wang, Zhengkai Jiang +4
cs.CVarXiv:2007.03496v32020A Simple Framework for Open-Vocabulary Segmentation and Detection
Hao Zhang, Feng Li, Xueyan Zou +5
cs.CVarXiv:2303.08131v32023Dual-Stream Semantic Guidance with Prototype Anchor Calibration for Source-Fully-Free Adaptation of Vision-Language Models
Weiwei Xiang, Shun Peng, Guangyi Xiao +2
cs.CVarXiv:2608.28145v12026DeepfakeBench: A Comprehensive Benchmark of Deepfake Detection
Zhiyuan Yan, Yong Zhang, Xinhang Yuan +2
cs.CVarXiv:2307.01426v22023PersFormer: 3D Lane Detection via Perspective Transformer and the OpenLane Benchmark
Li Chen, Chonghao Sima, Yang Li +8
cs.CVarXiv:2203.11089v32022Gait Recognition in the Wild with Dense 3D Representations and A Benchmark
Jinkai Zheng, Xinchen Liu, Wu Liu +3
cs.CVarXiv:2204.02569v12022Web-Scale Training for Face Identification
Yaniv Taigman, Ming Yang, Marc'Aurelio Ranzato +1
cs.CVarXiv:1406.5266v22014Physically-Based Rendering for Indoor Scene Understanding Using Convolutional Neural Networks
Yinda Zhang, Shuran Song, Ersin Yumer +4
cs.CVarXiv:1612.07429v32016LucidDreamer: Domain-free Generation of 3D Gaussian Splatting Scenes
Jaeyoung Chung, Suyoung Lee, Hyeongjin Nam +2
cs.CVarXiv:2311.13384v22023Missing MRI Pulse Sequence Synthesis using Multi-Modal Generative Adversarial Network
Anmol Sharma, Ghassan Hamarneh
eess.IVcs.AIcs.CVarXiv:1904.12200v32019Universal Litmus Patterns: Revealing Backdoor Attacks in CNNs
Soheil Kolouri, Aniruddha Saha, Hamed Pirsiavash +1
cs.CVarXiv:1906.10842v22019Geometry-Consistent Generative Adversarial Networks for One-Sided Unsupervised Domain Mapping
Huan Fu, Mingming Gong, Chaohui Wang +3
cs.CVarXiv:1809.05852v22018D3S -- A Discriminative Single Shot Segmentation Tracker
Alan Lukežič, Jiří Matas, Matej Kristan
cs.CVarXiv:1911.08862v22019Dual Residual Networks Leveraging the Potential of Paired Operations for Image Restoration
Xing Liu, Masanori Suganuma, Zhun Sun +1
cs.CVarXiv:1903.08817v22019A Unified Objective for Novel Class Discovery
Enrico Fini, Enver Sangineto, Stéphane Lathuilière +3
cs.CVcs.LGarXiv:2108.08536v42021Learning Video Representations from Large Language Models
Yue Zhao, Ishan Misra, Philipp Krähenbühl +1
cs.CVarXiv:2212.04501v12022A Comprehensive Review of Computer-aided Whole-slide Image Analysis: from Datasets to Feature Extraction, Segmentation, Classification, and Detection Approaches
Chen Li, Xintong Li, Md Rahaman +8
cs.CVcs.AIarXiv:2102.10553v12021