Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
12,361 to 12,420 of 18,839
Human-Aware Motion Deblurring
Ziyi Shen, Wenguan Wang, Xiankai Lu +4
cs.CVarXiv:2001.06816v12020RoNIN: Robust Neural Inertial Navigation in the Wild: Benchmark, Evaluations, and New Methods
Hang Yan, Sachini Herath, Yasutaka Furukawa
cs.CVcs.ROarXiv:1905.12853v12019Deep Image Deblurring: A Survey
Kaihao Zhang, Wenqi Ren, Wenhan Luo +4
cs.CVarXiv:2201.10700v22022A Survey on Hallucination in Large Vision-Language Models
Hanchao Liu, Wenyuan Xue, Yifei Chen +6
cs.CVcs.CLcs.LGarXiv:2402.00253v22024Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision
Long Cui, Xiaoqian Liu, Qi Qin +4
cs.CVarXiv:2608.16812v12026KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill
Yunxin Li, Jinchao Li, Shibo Su +7
cs.CLcs.CVarXiv:2607.12625v22026Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models
Mohammad Mahdi Abootorabi, Omid Ghahroodi, Anas Madkoor +4
cs.CVcs.AIcs.CLarXiv:2606.05531v12026Edge Intelligence: On-Demand Deep Learning Model Co-Inference with Device-Edge Synergy
En Li, Zhi Zhou, Xu Chen
cs.DCcs.AIcs.CVarXiv:1806.07840v42018Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation
Shuhong Zheng, Aashish Kumar Misraa, Yu-Teng Li +2
cs.CVcs.AIcs.GRarXiv:2605.26111v12026GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation
Sixiang Chen, Zhaohu Xing, Tian Ye +7
cs.CVarXiv:2605.21605v22026A$^2$RD: Agentic Autoregressive Diffusion for Long Video Consistency
Do Xuan Long, Yale Song, Min-Yen Kan +2
cs.CVcs.AIarXiv:2605.06924v12026Object-centric Auto-encoders and Dummy Anomalies for Abnormal Event Detection in Video
Radu Tudor Ionescu, Fahad Shahbaz Khan, Mariana-Iuliana Georgescu +1
cs.CVcs.LGarXiv:1812.04960v22018Accurate Leukocyte Detection Based on Deformable-DETR and Multi-Level Feature Fusion for Aiding Diagnosis of Blood Diseases
Yifei Chen, Chenyan Zhang, Ben Chen +8
cs.CVcs.AIarXiv:2401.00926v42024DatasetGAN: Efficient Labeled Data Factory with Minimal Human Effort
Yuxuan Zhang, Huan Ling, Jun Gao +5
cs.CVarXiv:2104.06490v22021DLOW: Domain Flow for Adaptation and Generalization
Rui Gong, Wen Li, Yuhua Chen +1
cs.CVarXiv:1812.05418v22018Image Matching Using SIFT, SURF, BRIEF and ORB: Performance Comparison for Distorted Images
Ebrahim Karami, Siva Prasad, Mohamed Shehata
cs.CVarXiv:1710.02726v12017CNN-Based Projected Gradient Descent for Consistent Image Reconstruction
Harshit Gupta, Kyong Hwan Jin, Ha Q. Nguyen +2
cs.CVarXiv:1709.01809v12017Structured Pruning for Deep Convolutional Neural Networks: A survey
Yang He, Lingao Xiao
cs.CVarXiv:2303.00566v22023Multi-Scale Vision Longformer: A New Vision Transformer for High-Resolution Image Encoding
Pengchuan Zhang, Xiyang Dai, Jianwei Yang +4
cs.CVcs.AIcs.LGarXiv:2103.15358v22021Dual Aggregation Transformer for Image Super-Resolution
Zheng Chen, Yulun Zhang, Jinjin Gu +3
cs.CVarXiv:2308.03364v22023CrowdNet: A Deep Convolutional Network for Dense Crowd Counting
Lokesh Boominathan, Srinivas S S Kruthiventi, R. Venkatesh Babu
cs.CVarXiv:1608.06197v12016UnitBox: An Advanced Object Detection Network
Jiahui Yu, Yuning Jiang, Zhangyang Wang +2
cs.CVarXiv:1608.01471v12016Training Deep Networks for Facial Expression Recognition with Crowd-Sourced Label Distribution
Emad Barsoum, Cha Zhang, Cristian Canton Ferrer +1
cs.CVarXiv:1608.01041v22016Multi-Scale Convolutional Neural Networks for Time Series Classification
Zhicheng Cui, Wenlin Chen, Yixin Chen
cs.CVarXiv:1603.06995v42016Vision-Language Pre-Training with Triple Contrastive Learning
Jinyu Yang, Jiali Duan, Son Tran +6
cs.CVarXiv:2202.10401v42022Lending Orientation to Neural Networks for Cross-view Geo-localization
Liu Liu, Hongdong Li
cs.CVarXiv:1903.12351v12019NoMaD: Goal Masked Diffusion Policies for Navigation and Exploration
Ajay Sridhar, Dhruv Shah, Catherine Glossop +1
cs.ROcs.CVcs.LGarXiv:2310.07896v12023Images Speak in Images: A Generalist Painter for In-Context Visual Learning
Xinlong Wang, Wen Wang, Yue Cao +2
cs.CVarXiv:2212.02499v22022Composer: Creative and Controllable Image Synthesis with Composable Conditions
Lianghua Huang, Di Chen, Yu Liu +3
cs.CVcs.GRarXiv:2302.09778v22023Class-Incremental Learning: A Survey
Da-Wei Zhou, Qi-Wei Wang, Zhi-Hong Qi +3
cs.CVcs.LGarXiv:2302.03648v22023Local Spectral Graph Convolution for Point Set Feature Learning
Chu Wang, Babak Samari, Kaleem Siddiqi
cs.CVcs.LGarXiv:1803.05827v12018Why rankings of biomedical image analysis competitions should be interpreted with care
Lena Maier-Hein, Matthias Eisenmann, Annika Reinke +35
cs.CVarXiv:1806.02051v22018CODE: Cross-Modal Calibration and Dynamic Suppression for Open World Object Detection
Hao Xu, Zhaoning Shi, Hehe Jin +1
cs.CVarXiv:2608.27214v12026Scaling & Shifting Your Features: A New Baseline for Efficient Model Tuning
Dongze Lian, Daquan Zhou, Jiashi Feng +1
cs.CVarXiv:2210.08823v32022StereoNet: Guided Hierarchical Refinement for Real-Time Edge-Aware Depth Prediction
Sameh Khamis, Sean Fanello, Christoph Rhemann +3
cs.CVarXiv:1807.08865v12018Revisiting Active Perception
Ruzena Bajcsy, Yiannis Aloimonos, John K. Tsotsos
cs.CVcs.ROarXiv:1603.02729v22016DINOcular: Self-Supervised Visuospatial Representations
Farkhat Almukhamedov, Sami Azirar, Hermann Blum
cs.CVarXiv:2608.27226v12026Gabor Convolutional Networks
Shangzhen Luan, Baochang Zhang, Chen Chen +3
cs.CVarXiv:1705.01450v42017FAN-LoRA: A Fourier-Adaptive Nonlinear Low-Rank Adaptor for Medical Foundation Model Domain Adaptation
Ziquan Liu, Zhewei Zhu, Xuyang Shi
cs.CVarXiv:2608.26531v12026Tips and Tricks for Visual Question Answering: Learnings from the 2017 Challenge
Damien Teney, Peter Anderson, Xiaodong He +1
cs.CVcs.CLarXiv:1708.02711v12017Convolutional Image Captioning
Jyoti Aneja, Aditya Deshpande, Alexander Schwing
cs.CVarXiv:1711.09151v12017OccFormer: Dual-path Transformer for Vision-based 3D Semantic Occupancy Prediction
Yunpeng Zhang, Zheng Zhu, Dalong Du
cs.CVarXiv:2304.05316v12023Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with Transformers
Zhaoshuo Li, Xingtong Liu, Nathan Drenkow +4
cs.CVarXiv:2011.02910v42020Unsupervised Adaptation of 3D CT Foundation Models for 3D CBCT Segmentation
Gauthier Miralles, Loic Le Folgoc, Vincent Jugnon +1
cs.CVarXiv:2608.27190v12026Learning to Generate Long-term Future via Hierarchical Prediction
Ruben Villegas, Jimei Yang, Yuliang Zou +3
cs.CVarXiv:1704.05831v52017Know Your Surroundings: Exploiting Scene Information for Object Tracking
Goutam Bhat, Martin Danelljan, Luc Van Gool +1
cs.CVarXiv:2003.11014v22020TokenPose: Learning Keypoint Tokens for Human Pose Estimation
Yanjie Li, Shoukui Zhang, Zhicheng Wang +4
cs.CVarXiv:2104.03516v32021Cross-Modality Fusion Transformer for Multispectral Object Detection
Fang Qingyun, Han Dapeng, Wang Zhaokui
eess.IVcs.CVarXiv:2111.00273v42021Learning to Predict Indoor Illumination from a Single Image
Marc-André Gardner, Kalyan Sunkavalli, Ersin Yumer +4
cs.CVcs.GRstat.MLarXiv:1704.00090v32017StyleSDF: High-Resolution 3D-Consistent Image and Geometry Generation
Roy Or-El, Xuan Luo, Mengyi Shan +3
cs.CVcs.AIcs.GRarXiv:2112.11427v22021IMRAM: Iterative Matching with Recurrent Attention Memory for Cross-Modal Image-Text Retrieval
Hui Chen, Guiguang Ding, Xudong Liu +3
cs.CVarXiv:2003.03772v12020Efficient Frequency Domain-based Transformers for High-Quality Image Deblurring
Lingshun Kong, Jiangxin Dong, Mingqiang Li +2
cs.CVarXiv:2211.12250v12022Focal and Global Knowledge Distillation for Detectors
Zhendong Yang, Zhe Li, Xiaohu Jiang +4
cs.CVarXiv:2111.11837v22021Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents
Boyu Gou, Ruohan Wang, Boyuan Zheng +5
cs.AIcs.CLcs.CVarXiv:2410.05243v32024AFPN: Asymptotic Feature Pyramid Network for Object Detection
Guoyu Yang, Jie Lei, Zhikuan Zhu +3
cs.CVarXiv:2306.15988v22023Coherent Semantic Attention for Image Inpainting
Hongyu Liu, Bin Jiang, Yi Xiao +1
cs.CVarXiv:1905.12384v32019Tracking Emerges by Colorizing Videos
Carl Vondrick, Abhinav Shrivastava, Alireza Fathi +2
cs.CVcs.GRcs.LGarXiv:1806.09594v22018LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
Yuzhang Shang, Mu Cai, Bingxin Xu +2
cs.CVcs.AIcs.CLarXiv:2403.15388v62024Convolutional Kernel Networks
Julien Mairal, Piotr Koniusz, Zaid Harchaoui +1
cs.CVcs.LGstat.MLarXiv:1406.3332v22014A Generative Adversarial Approach for Zero-Shot Learning from Noisy Texts
Yizhe Zhu, Mohamed Elhoseiny, Bingchen Liu +2
cs.CVarXiv:1712.01381v32017