Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

14,281 to 14,340 of 18,815

  1. Learning to Compose Dynamic Tree Structures for Visual Contexts

    Kaihua Tang, Hanwang Zhang, Baoyuan Wu +2

    cs.CVarXiv:1812.01880v12018
  2. $A^2$-Nets: Double Attention Networks

    Yunpeng Chen, Yannis Kalantidis, Jianshu Li +2

    cs.CVarXiv:1810.11579v12018
  3. A study of the effect of JPG compression on adversarial images

    Gintare Karolina Dziugaite, Zoubin Ghahramani, Daniel M. Roy

    cs.CVcs.LGarXiv:1608.00853v12016
  4. Asymmetric Tri-training for Unsupervised Domain Adaptation

    Kuniaki Saito, Yoshitaka Ushiku, Tatsuya Harada

    cs.CVcs.AIarXiv:1702.08400v32017
  5. DenseTNT: End-to-end Trajectory Prediction from Dense Goal Sets

    Junru Gu, Chen Sun, Hang Zhao

    cs.CVcs.ROarXiv:2108.09640v22021
  6. DDMS: Discriminative Distillation of Multi-view Foundational Features into Single-view Models

    Jeong-gi Kwak, Sho Kagami, Yuki Ono +1

    cs.CVarXiv:2608.23850v12026
  7. MirrorGAN: Learning Text-to-image Generation by Redescription

    Tingting Qiao, Jing Zhang, Duanqing Xu +1

    cs.CLcs.CVcs.LGarXiv:1903.05854v12019
  8. CNN Image Retrieval Learns from BoW: Unsupervised Fine-Tuning with Hard Examples

    Filip Radenović, Giorgos Tolias, Ondřej Chum

    cs.CVarXiv:1604.02426v32016
  9. Too much of a good thing -- when knowledge distillation promotes overfitting, and how to avoid it

    Irene Trigueros-Lorca, Leonardo Concepción, Christian Wagner +2

    cs.CVcs.AIarXiv:2608.23752v12026
  10. Learning to Poke by Poking: Experiential Learning of Intuitive Physics

    Pulkit Agrawal, Ashvin Nair, Pieter Abbeel +2

    cs.CVcs.AIcs.ROarXiv:1606.07419v22016
  11. Mapping the Concept Landscape: Structural Perception of Global Distributions for Transparent Data Pruning

    Dongyue Wu, Tao Ma

    cs.LGcs.CVarXiv:2608.22858v12026
  12. Multi-Stage Prompt-Guided Feature Modulation for Generalizable Brain Tumor Segmentation

    Mohammad Mahdi Danesh Pajouh, Sara Saeedi

    eess.IVcs.CVarXiv:2608.23745v12026
  13. 3D Terrestrial lidar data classification of complex natural scenes using a multi-scale dimensionality criterion: applications in geomorphology

    Nicolas Brodu, Dimitri Lague

    cs.CVphysics.geo-pharXiv:1107.0550v32011
  14. Deep Colorization

    Zezhou Cheng, Qingxiong Yang, Bin Sheng

    cs.CVarXiv:1605.00075v12016
  15. MotionCtrl: A Unified and Flexible Motion Controller for Video Generation

    Zhouxia Wang, Ziyang Yuan, Xintao Wang +4

    cs.CVcs.AIcs.LGarXiv:2312.03641v22023
  16. DepthTransfer: Depth Extraction from Video Using Non-parametric Sampling

    Kevin Karsch, Ce Liu, Sing Bing Kang

    cs.CVarXiv:2001.00987v12019
  17. TONAV: Task-Oriented Navigation and Action-Velocity Chunk Learning for Articulated Object Quadrupedal Mobile Manipulation

    Haoran Lin, Mingyu Yang, Pengfei Qi +6

    cs.ROcs.CVarXiv:2608.22296v12026
  18. Latent-NeRF for Shape-Guided Generation of 3D Shapes and Textures

    Gal Metzer, Elad Richardson, Or Patashnik +2

    cs.CVcs.GRarXiv:2211.07600v12022
  19. Using Trusted Data to Train Deep Networks on Labels Corrupted by Severe Noise

    Dan Hendrycks, Mantas Mazeika, Duncan Wilson +1

    cs.LGcs.CLcs.CVarXiv:1802.05300v42018
  20. An Empirical Study and Analysis of Generalized Zero-Shot Learning for Object Recognition in the Wild

    Wei-Lun Chao, Soravit Changpinyo, Boqing Gong +1

    cs.CVarXiv:1605.04253v22016
  21. Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments

    Jacob Krantz, Erik Wijmans, Arjun Majumdar +2

    cs.CVcs.CLcs.ROarXiv:2004.02857v22020
  22. Dataset Distillation by Matching Training Trajectories

    George Cazenavette, Tongzhou Wang, Antonio Torralba +2

    cs.CVcs.AIcs.LGarXiv:2203.11932v12022
  23. A Normalized Gaussian Wasserstein Distance for Tiny Object Detection

    Jinwang Wang, Chang Xu, Wen Yang +1

    cs.CVarXiv:2110.13389v22021
  24. Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model

    Ruoxi Shi, Hansheng Chen, Zhuoyang Zhang +6

    cs.CVcs.GRarXiv:2310.15110v12023
  25. Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulation

    Xiwen Chen, Zelin Li, Zhiruo Zhou +3

    cs.ROcs.AIcs.CVarXiv:2608.23138v12026
  26. Minimally distorted Adversarial Examples with a Fast Adaptive Boundary Attack

    Francesco Croce, Matthias Hein

    cs.LGcs.CRcs.CVarXiv:1907.02044v22019
  27. DISK: Learning local features with policy gradient

    Michał J. Tyszkiewicz, Pascal Fua, Eduard Trulls

    cs.CVcs.LGarXiv:2006.13566v22020
  28. View Adaptive Recurrent Neural Networks for High Performance Human Action Recognition from Skeleton Data

    Pengfei Zhang, Cuiling Lan, Junliang Xing +3

    cs.CVarXiv:1703.08274v22017
  29. NeoWorld-Pro: Programming Interactive Scenes from Monocular Images for Embodied Simulation

    Yumeng He, Yichen Song, Xiaotian Yang +5

    cs.CVarXiv:2608.24212v12026
  30. UniFormer: Unifying Convolution and Self-attention for Visual Recognition

    Kunchang Li, Yali Wang, Junhao Zhang +5

    cs.CVarXiv:2201.09450v32022
  31. TransPhy: Visual In-Context Learning for Physically Grounded Image Editing

    Siyi Xie, Xuanke Shi, Jinsheng Quan +4

    cs.CVcs.AIarXiv:2608.24119v12026
  32. Why do deep convolutional networks generalize so poorly to small image transformations?

    Aharon Azulay, Yair Weiss

    cs.CVarXiv:1805.12177v42018
  33. MONet: Unsupervised Scene Decomposition and Representation

    Christopher P. Burgess, Loic Matthey, Nicholas Watters +4

    cs.CVcs.LGstat.MLarXiv:1901.11390v12019
  34. ExMesh++: From Multi-View Images to Relightable UV-PBR Mesh Assets via Topology-Adaptive Reconstruction and Decomposition

    Chuanjin Fan, Lifan Wu, Wenjie Chang +3

    cs.GRcs.CVarXiv:2608.24109v12026
  35. Data-Driven Sparse Structure Selection for Deep Neural Networks

    Zehao Huang, Naiyan Wang

    cs.CVcs.LGcs.NEarXiv:1707.01213v32017
  36. ViSculpt: Visual-Centric Agentic Geometry Editing

    Bo Pang, Jiaqi Pan, Xiaocheng Zhang +3

    cs.CVcs.GRcs.HCarXiv:2608.24169v12026
  37. A review: Deep learning for medical image segmentation using multi-modality fusion

    Tongxue Zhou, Su Ruan, Stéphane Canu

    eess.IVcs.CVcs.LGarXiv:2004.10664v22020
  38. Reflection Backdoor: A Natural Backdoor Attack on Deep Neural Networks

    Yunfei Liu, Xingjun Ma, James Bailey +1

    cs.CVarXiv:2007.02343v22020
  39. Adversarial Learning for Semi-Supervised Semantic Segmentation

    Wei-Chih Hung, Yi-Hsuan Tsai, Yan-Ting Liou +2

    cs.CVarXiv:1802.07934v22018
  40. Diffusion Autoencoders: Toward a Meaningful and Decodable Representation

    Konpat Preechakul, Nattanat Chatthee, Suttisak Wizadwongsa +1

    cs.CVcs.LGarXiv:2111.15640v32021
  41. Improved Distribution Matching Distillation for Fast Image Synthesis

    Tianwei Yin, Michaël Gharbi, Taesung Park +4

    cs.CVarXiv:2405.14867v22024
  42. See through Gradients: Image Batch Recovery via GradInversion

    Hongxu Yin, Arun Mallya, Arash Vahdat +3

    cs.LGcs.CVarXiv:2104.07586v12021
  43. Detecting Deepfakes with Self-Blended Images

    Kaede Shiohara, Toshihiko Yamasaki

    cs.CVarXiv:2204.08376v12022
  44. Deep Learning in Medical Image Registration: A Review

    Yabo Fu, Yang Lei, Tonghe Wang +3

    eess.IVcs.CVcs.LGarXiv:1912.12318v12019
  45. A Survey on Active Learning and Human-in-the-Loop Deep Learning for Medical Image Analysis

    Samuel Budd, Emma C Robinson, Bernhard Kainz

    cs.LGcs.CVcs.HCarXiv:1910.02923v22019
  46. DISN: Deep Implicit Surface Network for High-quality Single-view 3D Reconstruction

    Qiangeng Xu, Weiyue Wang, Duygu Ceylan +2

    cs.CVarXiv:1905.10711v52019
  47. E2S-Pruner: Progressive Two-Stage Evidence Fusion for Visual Token Pruning in Vision-Language Models

    Taoyu Qian, Qi Wang, Daqian Shi +3

    cs.CVcs.AIarXiv:2608.23253v12026
  48. Large Selective Kernel Network for Remote Sensing Object Detection

    Yuxuan Li, Qibin Hou, Zhaohui Zheng +3

    cs.CVarXiv:2303.09030v22023
  49. Adversarial Complementary Learning for Weakly Supervised Object Localization

    Xiaolin Zhang, Yunchao Wei, Jiashi Feng +2

    cs.CVarXiv:1804.06962v12018
  50. LRS3-TED: a large-scale dataset for visual speech recognition

    Triantafyllos Afouras, Joon Son Chung, Andrew Zisserman

    cs.CVarXiv:1809.00496v22018
  51. PredRNN++: Towards A Resolution of the Deep-in-Time Dilemma in Spatiotemporal Predictive Learning

    Yunbo Wang, Zhifeng Gao, Mingsheng Long +2

    cs.LGcs.CVstat.MLarXiv:1804.06300v22018
  52. Grad-CAM: Why did you say that?

    Ramprasaath R Selvaraju, Abhishek Das, Ramakrishna Vedantam +3

    stat.MLcs.CVcs.LGarXiv:1611.07450v22016
  53. One-Shot Visual Imitation Learning via Meta-Learning

    Chelsea Finn, Tianhe Yu, Tianhao Zhang +2

    cs.LGcs.AIcs.CVarXiv:1709.04905v12017
  54. AffineTok: Semantic Affine Consistency for Diffusion-Friendly Visual Tokenizer

    Junqiu Yu, Pandeng Li, Yikai Wang +11

    cs.CVarXiv:2608.23864v12026
  55. INTERACTION Dataset: An INTERnational, Adversarial and Cooperative moTION Dataset in Interactive Driving Scenarios with Semantic Maps

    Wei Zhan, Liting Sun, Di Wang +8

    cs.ROcs.CVcs.LGarXiv:1910.03088v12019
  56. Event-based Vision meets Deep Learning on Steering Prediction for Self-driving Cars

    Ana I. Maqueda, Antonio Loquercio, Guillermo Gallego +2

    cs.CVcs.LGcs.ROarXiv:1804.01310v12018
  57. Open-Vocabulary Object Detection Using Captions

    Alireza Zareian, Kevin Dela Rosa, Derek Hao Hu +1

    cs.CVcs.AIcs.LGarXiv:2011.10678v22020
  58. Think Only When Needed: Prompt-Authority Control for Selective Slow-Path Intervention in Vision-Language-Action Manipulation

    Zhiruo Zhou, Zelin Li, Xiwen Chen +4

    cs.ROcs.AIcs.CVarXiv:2608.23224v12026
  59. MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

    Zhengyuan Yang, Linjie Li, Jianfeng Wang +7

    cs.CVcs.CLcs.LGarXiv:2303.11381v12023
  60. Erasing Concepts from Diffusion Models

    Rohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman +1

    cs.CVarXiv:2303.07345v32023