Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
14,281 to 14,340 of 18,815
Learning to Compose Dynamic Tree Structures for Visual Contexts
Kaihua Tang, Hanwang Zhang, Baoyuan Wu +2
cs.CVarXiv:1812.01880v12018$A^2$-Nets: Double Attention Networks
Yunpeng Chen, Yannis Kalantidis, Jianshu Li +2
cs.CVarXiv:1810.11579v12018A study of the effect of JPG compression on adversarial images
Gintare Karolina Dziugaite, Zoubin Ghahramani, Daniel M. Roy
cs.CVcs.LGarXiv:1608.00853v12016Asymmetric Tri-training for Unsupervised Domain Adaptation
Kuniaki Saito, Yoshitaka Ushiku, Tatsuya Harada
cs.CVcs.AIarXiv:1702.08400v32017DenseTNT: End-to-end Trajectory Prediction from Dense Goal Sets
Junru Gu, Chen Sun, Hang Zhao
cs.CVcs.ROarXiv:2108.09640v22021DDMS: Discriminative Distillation of Multi-view Foundational Features into Single-view Models
Jeong-gi Kwak, Sho Kagami, Yuki Ono +1
cs.CVarXiv:2608.23850v12026MirrorGAN: Learning Text-to-image Generation by Redescription
Tingting Qiao, Jing Zhang, Duanqing Xu +1
cs.CLcs.CVcs.LGarXiv:1903.05854v12019CNN Image Retrieval Learns from BoW: Unsupervised Fine-Tuning with Hard Examples
Filip Radenović, Giorgos Tolias, Ondřej Chum
cs.CVarXiv:1604.02426v32016Too much of a good thing -- when knowledge distillation promotes overfitting, and how to avoid it
Irene Trigueros-Lorca, Leonardo Concepción, Christian Wagner +2
cs.CVcs.AIarXiv:2608.23752v12026Learning to Poke by Poking: Experiential Learning of Intuitive Physics
Pulkit Agrawal, Ashvin Nair, Pieter Abbeel +2
cs.CVcs.AIcs.ROarXiv:1606.07419v22016Mapping the Concept Landscape: Structural Perception of Global Distributions for Transparent Data Pruning
Dongyue Wu, Tao Ma
cs.LGcs.CVarXiv:2608.22858v12026Multi-Stage Prompt-Guided Feature Modulation for Generalizable Brain Tumor Segmentation
Mohammad Mahdi Danesh Pajouh, Sara Saeedi
eess.IVcs.CVarXiv:2608.23745v120263D Terrestrial lidar data classification of complex natural scenes using a multi-scale dimensionality criterion: applications in geomorphology
Nicolas Brodu, Dimitri Lague
cs.CVphysics.geo-pharXiv:1107.0550v32011Deep Colorization
Zezhou Cheng, Qingxiong Yang, Bin Sheng
cs.CVarXiv:1605.00075v12016MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
Zhouxia Wang, Ziyang Yuan, Xintao Wang +4
cs.CVcs.AIcs.LGarXiv:2312.03641v22023DepthTransfer: Depth Extraction from Video Using Non-parametric Sampling
Kevin Karsch, Ce Liu, Sing Bing Kang
cs.CVarXiv:2001.00987v12019TONAV: Task-Oriented Navigation and Action-Velocity Chunk Learning for Articulated Object Quadrupedal Mobile Manipulation
Haoran Lin, Mingyu Yang, Pengfei Qi +6
cs.ROcs.CVarXiv:2608.22296v12026Latent-NeRF for Shape-Guided Generation of 3D Shapes and Textures
Gal Metzer, Elad Richardson, Or Patashnik +2
cs.CVcs.GRarXiv:2211.07600v12022Using Trusted Data to Train Deep Networks on Labels Corrupted by Severe Noise
Dan Hendrycks, Mantas Mazeika, Duncan Wilson +1
cs.LGcs.CLcs.CVarXiv:1802.05300v42018An Empirical Study and Analysis of Generalized Zero-Shot Learning for Object Recognition in the Wild
Wei-Lun Chao, Soravit Changpinyo, Boqing Gong +1
cs.CVarXiv:1605.04253v22016Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments
Jacob Krantz, Erik Wijmans, Arjun Majumdar +2
cs.CVcs.CLcs.ROarXiv:2004.02857v22020Dataset Distillation by Matching Training Trajectories
George Cazenavette, Tongzhou Wang, Antonio Torralba +2
cs.CVcs.AIcs.LGarXiv:2203.11932v12022A Normalized Gaussian Wasserstein Distance for Tiny Object Detection
Jinwang Wang, Chang Xu, Wen Yang +1
cs.CVarXiv:2110.13389v22021Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model
Ruoxi Shi, Hansheng Chen, Zhuoyang Zhang +6
cs.CVcs.GRarXiv:2310.15110v12023Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulation
Xiwen Chen, Zelin Li, Zhiruo Zhou +3
cs.ROcs.AIcs.CVarXiv:2608.23138v12026Minimally distorted Adversarial Examples with a Fast Adaptive Boundary Attack
Francesco Croce, Matthias Hein
cs.LGcs.CRcs.CVarXiv:1907.02044v22019DISK: Learning local features with policy gradient
Michał J. Tyszkiewicz, Pascal Fua, Eduard Trulls
cs.CVcs.LGarXiv:2006.13566v22020View Adaptive Recurrent Neural Networks for High Performance Human Action Recognition from Skeleton Data
Pengfei Zhang, Cuiling Lan, Junliang Xing +3
cs.CVarXiv:1703.08274v22017NeoWorld-Pro: Programming Interactive Scenes from Monocular Images for Embodied Simulation
Yumeng He, Yichen Song, Xiaotian Yang +5
cs.CVarXiv:2608.24212v12026UniFormer: Unifying Convolution and Self-attention for Visual Recognition
Kunchang Li, Yali Wang, Junhao Zhang +5
cs.CVarXiv:2201.09450v32022TransPhy: Visual In-Context Learning for Physically Grounded Image Editing
Siyi Xie, Xuanke Shi, Jinsheng Quan +4
cs.CVcs.AIarXiv:2608.24119v12026Why do deep convolutional networks generalize so poorly to small image transformations?
Aharon Azulay, Yair Weiss
cs.CVarXiv:1805.12177v42018MONet: Unsupervised Scene Decomposition and Representation
Christopher P. Burgess, Loic Matthey, Nicholas Watters +4
cs.CVcs.LGstat.MLarXiv:1901.11390v12019ExMesh++: From Multi-View Images to Relightable UV-PBR Mesh Assets via Topology-Adaptive Reconstruction and Decomposition
Chuanjin Fan, Lifan Wu, Wenjie Chang +3
cs.GRcs.CVarXiv:2608.24109v12026Data-Driven Sparse Structure Selection for Deep Neural Networks
Zehao Huang, Naiyan Wang
cs.CVcs.LGcs.NEarXiv:1707.01213v32017ViSculpt: Visual-Centric Agentic Geometry Editing
Bo Pang, Jiaqi Pan, Xiaocheng Zhang +3
cs.CVcs.GRcs.HCarXiv:2608.24169v12026A review: Deep learning for medical image segmentation using multi-modality fusion
Tongxue Zhou, Su Ruan, Stéphane Canu
eess.IVcs.CVcs.LGarXiv:2004.10664v22020Reflection Backdoor: A Natural Backdoor Attack on Deep Neural Networks
Yunfei Liu, Xingjun Ma, James Bailey +1
cs.CVarXiv:2007.02343v22020Adversarial Learning for Semi-Supervised Semantic Segmentation
Wei-Chih Hung, Yi-Hsuan Tsai, Yan-Ting Liou +2
cs.CVarXiv:1802.07934v22018Diffusion Autoencoders: Toward a Meaningful and Decodable Representation
Konpat Preechakul, Nattanat Chatthee, Suttisak Wizadwongsa +1
cs.CVcs.LGarXiv:2111.15640v32021Improved Distribution Matching Distillation for Fast Image Synthesis
Tianwei Yin, Michaël Gharbi, Taesung Park +4
cs.CVarXiv:2405.14867v22024See through Gradients: Image Batch Recovery via GradInversion
Hongxu Yin, Arun Mallya, Arash Vahdat +3
cs.LGcs.CVarXiv:2104.07586v12021Detecting Deepfakes with Self-Blended Images
Kaede Shiohara, Toshihiko Yamasaki
cs.CVarXiv:2204.08376v12022Deep Learning in Medical Image Registration: A Review
Yabo Fu, Yang Lei, Tonghe Wang +3
eess.IVcs.CVcs.LGarXiv:1912.12318v12019A Survey on Active Learning and Human-in-the-Loop Deep Learning for Medical Image Analysis
Samuel Budd, Emma C Robinson, Bernhard Kainz
cs.LGcs.CVcs.HCarXiv:1910.02923v22019DISN: Deep Implicit Surface Network for High-quality Single-view 3D Reconstruction
Qiangeng Xu, Weiyue Wang, Duygu Ceylan +2
cs.CVarXiv:1905.10711v52019E2S-Pruner: Progressive Two-Stage Evidence Fusion for Visual Token Pruning in Vision-Language Models
Taoyu Qian, Qi Wang, Daqian Shi +3
cs.CVcs.AIarXiv:2608.23253v12026Large Selective Kernel Network for Remote Sensing Object Detection
Yuxuan Li, Qibin Hou, Zhaohui Zheng +3
cs.CVarXiv:2303.09030v22023Adversarial Complementary Learning for Weakly Supervised Object Localization
Xiaolin Zhang, Yunchao Wei, Jiashi Feng +2
cs.CVarXiv:1804.06962v12018LRS3-TED: a large-scale dataset for visual speech recognition
Triantafyllos Afouras, Joon Son Chung, Andrew Zisserman
cs.CVarXiv:1809.00496v22018PredRNN++: Towards A Resolution of the Deep-in-Time Dilemma in Spatiotemporal Predictive Learning
Yunbo Wang, Zhifeng Gao, Mingsheng Long +2
cs.LGcs.CVstat.MLarXiv:1804.06300v22018Grad-CAM: Why did you say that?
Ramprasaath R Selvaraju, Abhishek Das, Ramakrishna Vedantam +3
stat.MLcs.CVcs.LGarXiv:1611.07450v22016One-Shot Visual Imitation Learning via Meta-Learning
Chelsea Finn, Tianhe Yu, Tianhao Zhang +2
cs.LGcs.AIcs.CVarXiv:1709.04905v12017AffineTok: Semantic Affine Consistency for Diffusion-Friendly Visual Tokenizer
Junqiu Yu, Pandeng Li, Yikai Wang +11
cs.CVarXiv:2608.23864v12026INTERACTION Dataset: An INTERnational, Adversarial and Cooperative moTION Dataset in Interactive Driving Scenarios with Semantic Maps
Wei Zhan, Liting Sun, Di Wang +8
cs.ROcs.CVcs.LGarXiv:1910.03088v12019Event-based Vision meets Deep Learning on Steering Prediction for Self-driving Cars
Ana I. Maqueda, Antonio Loquercio, Guillermo Gallego +2
cs.CVcs.LGcs.ROarXiv:1804.01310v12018Open-Vocabulary Object Detection Using Captions
Alireza Zareian, Kevin Dela Rosa, Derek Hao Hu +1
cs.CVcs.AIcs.LGarXiv:2011.10678v22020Think Only When Needed: Prompt-Authority Control for Selective Slow-Path Intervention in Vision-Language-Action Manipulation
Zhiruo Zhou, Zelin Li, Xiwen Chen +4
cs.ROcs.AIcs.CVarXiv:2608.23224v12026MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
Zhengyuan Yang, Linjie Li, Jianfeng Wang +7
cs.CVcs.CLcs.LGarXiv:2303.11381v12023Erasing Concepts from Diffusion Models
Rohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman +1
cs.CVarXiv:2303.07345v32023