Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

12,001 to 12,060 of 18,821

  1. Multi-Scale Spatial Temporal Graph Convolutional Network for Skeleton-Based Action Recognition

    Zhan Chen, Sicheng Li, Bing Yang +2

    cs.CVarXiv:2206.13028v12022
  2. Yin and Yang: Balancing and Answering Binary Visual Questions

    Peng Zhang, Yash Goyal, Douglas Summers-Stay +2

    cs.CLcs.CVcs.LGarXiv:1511.05099v52015
  3. EvalCrafter: Benchmarking and Evaluating Large Video Generation Models

    Yaofang Liu, Xiaodong Cun, Xuebo Liu +7

    cs.CVarXiv:2310.11440v32023
  4. OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents

    Hugo Laurençon, Lucile Saulnier, Léo Tronchon +9

    cs.IRcs.CVarXiv:2306.16527v22023
  5. Earthformer: Exploring Space-Time Transformers for Earth System Forecasting

    Zhihan Gao, Xingjian Shi, Hao Wang +4

    cs.LGcs.AIcs.CVarXiv:2207.05833v22022
  6. ControlVideo: Training-free Controllable Text-to-Video Generation

    Yabo Zhang, Yuxiang Wei, Dongsheng Jiang +3

    cs.CVarXiv:2305.13077v12023
  7. SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery

    Xin Guo, Jiangwei Lao, Bo Dang +13

    cs.CVarXiv:2312.10115v22023
  8. Smart Mining for Deep Metric Learning

    Ben Harwood, Vijay Kumar B G, Gustavo Carneiro +2

    cs.CVarXiv:1704.01285v32017
  9. Video-P2P: Video Editing with Cross-attention Control

    Shaoteng Liu, Yuechen Zhang, Wenbo Li +2

    cs.CVarXiv:2303.04761v12023
  10. SGCN:Sparse Graph Convolution Network for Pedestrian Trajectory Prediction

    Liushuai Shi, Le Wang, Chengjiang Long +4

    cs.CVarXiv:2104.01528v12021
  11. Learning to diagnose from scratch by exploiting dependencies among labels

    Li Yao, Eric Poblenz, Dmitry Dagunts +3

    cs.CVarXiv:1710.10501v22017
  12. Keeping Your Eye on the Ball: Trajectory Attention in Video Transformers

    Mandela Patrick, Dylan Campbell, Yuki M. Asano +5

    cs.CVarXiv:2106.05392v22021
  13. What does a platypus look like? Generating customized prompts for zero-shot image classification

    Sarah Pratt, Ian Covert, Rosanne Liu +1

    cs.CVcs.LGarXiv:2209.03320v32022
  14. Adversarial Spatio-Temporal Learning for Video Deblurring

    Kaihao Zhang, Wenhan Luo, Yiran Zhong +3

    cs.CVarXiv:1804.00533v22018
  15. Sharp U-Net: Depthwise Convolutional Network for Biomedical Image Segmentation

    Hasib Zunair, A. Ben Hamza

    eess.IVcs.CVarXiv:2107.12461v12021
  16. Progressive Domain Adaptation for Object Detection

    Han-Kai Hsu, Chun-Han Yao, Yi-Hsuan Tsai +4

    cs.CVarXiv:1910.11319v12019
  17. VectorMapNet: End-to-end Vectorized HD Map Learning

    Yicheng Liu, Tianyuan Yuan, Yue Wang +2

    cs.CVcs.ROarXiv:2206.08920v62022
  18. Modeling the Background for Incremental Learning in Semantic Segmentation

    Fabio Cermelli, Massimiliano Mancini, Samuel Rota Bulò +2

    cs.CVarXiv:2002.00718v22020
  19. Flexible Diffusion Modeling of Long Videos

    William Harvey, Saeid Naderiparizi, Vaden Masrani +2

    cs.CVcs.LGarXiv:2205.11495v32022
  20. In-Place Activated BatchNorm for Memory-Optimized Training of DNNs

    Samuel Rota Bulò, Lorenzo Porzi, Peter Kontschieder

    cs.CVarXiv:1712.02616v32017
  21. Appearance-and-Relation Networks for Video Classification

    Limin Wang, Wei Li, Wen Li +1

    cs.CVarXiv:1711.09125v22017
  22. UniFormer: Unified Transformer for Efficient Spatiotemporal Representation Learning

    Kunchang Li, Yali Wang, Peng Gao +4

    cs.CVarXiv:2201.04676v32022
  23. Skeleton-Based Action Recognition with Spatial Reasoning and Temporal Stack Learning

    Chenyang Si, Ya Jing, Wei Wang +2

    cs.CVarXiv:1805.02335v22018
  24. Morphing and Sampling Network for Dense Point Cloud Completion

    Minghua Liu, Lu Sheng, Sheng Yang +2

    cs.CVarXiv:1912.00280v12019
  25. Building a Large Scale Dataset for Image Emotion Recognition: The Fine Print and The Benchmark

    Quanzeng You, Jiebo Luo, Hailin Jin +1

    cs.AIcs.CVarXiv:1605.02677v12016
  26. Residual Networks of Residual Networks: Multilevel Residual Networks

    Ke Zhang, Miao Sun, Tony X. Han +3

    cs.CVarXiv:1608.02908v22016
  27. Video Summarization with Attention-Based Encoder-Decoder Networks

    Zhong Ji, Kailin Xiong, Yanwei Pang +1

    cs.CVarXiv:1708.09545v22017
  28. Learning Two-View Correspondences and Geometry Using Order-Aware Network

    Jiahui Zhang, Dawei Sun, Zixin Luo +6

    cs.CVcs.CGcs.LGarXiv:1908.04964v12019
  29. Deep Learning Based Brain Tumor Segmentation: A Survey

    Zhihua Liu, Lei Tong, Zheheng Jiang +6

    eess.IVcs.CVarXiv:2007.09479v32020
  30. Self-supervised Learning in Remote Sensing: A Review

    Yi Wang, Conrad M Albrecht, Nassim Ait Ali Braham +2

    cs.CVarXiv:2206.13188v22022
  31. NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario

    Tianwen Qian, Jingjing Chen, Linhai Zhuo +2

    cs.CVarXiv:2305.14836v22023
  32. Neural Nearest Neighbors Networks

    Tobias Plötz, Stefan Roth

    cs.CVcs.LGarXiv:1810.12575v12018
  33. DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data

    Stephanie Fu, Netanel Tamir, Shobhita Sundaram +4

    cs.CVcs.LGarXiv:2306.09344v32023
  34. Camera Distance-aware Top-down Approach for 3D Multi-person Pose Estimation from a Single RGB Image

    Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee

    cs.CVarXiv:1907.11346v22019
  35. RoadTracer: Automatic Extraction of Road Networks from Aerial Images

    Favyen Bastani, Songtao He, Sofiane Abbar +5

    cs.CVarXiv:1802.03680v22018
  36. Classification of Hyperspectral and LiDAR Data Using Coupled CNNs

    Renlong Hang, Zhu Li, Pedram Ghamisi +3

    cs.CVeess.IVarXiv:2002.01144v12020
  37. Low-rank Bilinear Pooling for Fine-Grained Classification

    Shu Kong, Charless Fowlkes

    cs.CVarXiv:1611.05109v22016
  38. Compressed Video Action Recognition

    Chao-Yuan Wu, Manzil Zaheer, Hexiang Hu +3

    cs.CVarXiv:1712.00636v22017
  39. Neural Prototype Trees for Interpretable Fine-grained Image Recognition

    Meike Nauta, Ron van Bree, Christin Seifert

    cs.CVcs.AIcs.LGarXiv:2012.02046v22020
  40. Uncertainty-Aware Blind Image Quality Assessment in the Laboratory and Wild

    Weixia Zhang, Kede Ma, Guangtao Zhai +1

    cs.CVcs.LGcs.MMarXiv:2005.13983v62020
  41. Disentangled Non-Local Neural Networks

    Minghao Yin, Zhuliang Yao, Yue Cao +4

    cs.CVcs.CLcs.LGarXiv:2006.06668v22020
  42. How Neural Networks Extrapolate: From Feedforward to Graph Neural Networks

    Keyulu Xu, Mozhi Zhang, Jingling Li +3

    cs.LGcs.AIcs.CVarXiv:2009.11848v52020
  43. Stereo DSO: Large-Scale Direct Sparse Visual Odometry with Stereo Cameras

    Rui Wang, Martin Schwörer, Daniel Cremers

    cs.CVarXiv:1708.07878v12017
  44. Designing Deep Networks for Surface Normal Estimation

    Xiaolong Wang, David F. Fouhey, Abhinav Gupta

    cs.CVarXiv:1411.4958v12014
  45. Multi-task Collaborative Network for Joint Referring Expression Comprehension and Segmentation

    Gen Luo, Yiyi Zhou, Xiaoshuai Sun +4

    cs.CVarXiv:2003.08813v12020
  46. Appearance-Based Loop Closure Detection for Online Large-Scale and Long-Term Operation

    Mathieu Labbé, François Michaud

    cs.ROcs.CVarXiv:2407.15304v12024
  47. Dual Motion GAN for Future-Flow Embedded Video Prediction

    Xiaodan Liang, Lisa Lee, Wei Dai +1

    cs.CVarXiv:1708.00284v22017
  48. TOPIQ: A Top-down Approach from Semantics to Distortions for Image Quality Assessment

    Chaofeng Chen, Jiadi Mo, Jingwen Hou +5

    cs.CVarXiv:2308.03060v12023
  49. Plug-and-Play Priors for Bright Field Electron Tomography and Sparse Interpolation

    Suhas Sreehari, S. V. Venkatakrishnan, Brendt Wohlberg +3

    cs.CVeess.IVarXiv:1512.07331v12015
  50. Probabilistic Face Embeddings

    Yichun Shi, Anil K. Jain

    cs.CVarXiv:1904.09658v42019
  51. FaceScape: a Large-scale High Quality 3D Face Dataset and Detailed Riggable 3D Face Prediction

    Haotian Yang, Hao Zhu, Yanru Wang +4

    cs.CVarXiv:2003.13989v32020
  52. Hough-CNN: Deep Learning for Segmentation of Deep Brain Regions in MRI and Ultrasound

    Fausto Milletari, Seyed-Ahmad Ahmadi, Christine Kroll +8

    cs.CVarXiv:1601.07014v32016
  53. PolyGen: An Autoregressive Generative Model of 3D Meshes

    Charlie Nash, Yaroslav Ganin, S. M. Ali Eslami +1

    cs.GRcs.CVcs.LGarXiv:2002.10880v12020
  54. HP-GAN: Probabilistic 3D human motion prediction via GAN

    Emad Barsoum, John Kender, Zicheng Liu

    cs.CVcs.AIcs.HCarXiv:1711.09561v12017
  55. Variational Denoising Network: Toward Blind Noise Modeling and Removal

    Zongsheng Yue, Hongwei Yong, Qian Zhao +2

    cs.CVarXiv:1908.11314v52019
  56. Can Deep Learning Outperform Modern Commercial CT Image Reconstruction Methods?

    Hongming Shan, Atul Padole, Fatemeh Homayounieh +5

    cs.CVphysics.med-pharXiv:1811.03691v12018
  57. UC-Net: Uncertainty Inspired RGB-D Saliency Detection via Conditional Variational Autoencoders

    Jing Zhang, Deng-Ping Fan, Yuchao Dai +4

    cs.CVarXiv:2004.05763v12020
  58. Synergistic Image and Feature Adaptation: Towards Cross-Modality Domain Adaptation for Medical Image Segmentation

    Cheng Chen, Qi Dou, Hao Chen +2

    cs.CVarXiv:1901.08211v42019
  59. DecideNet: Counting Varying Density Crowds Through Attention Guided Detection and Density Estimation

    Jiang Liu, Chenqiang Gao, Deyu Meng +1

    cs.CVarXiv:1712.06679v22017
  60. YOLOv6 v3.0: A Full-Scale Reloading

    Chuyi Li, Lulu Li, Yifei Geng +6

    cs.CVarXiv:2301.05586v12023