Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
12,001 to 12,060 of 18,821
Multi-Scale Spatial Temporal Graph Convolutional Network for Skeleton-Based Action Recognition
Zhan Chen, Sicheng Li, Bing Yang +2
cs.CVarXiv:2206.13028v12022Yin and Yang: Balancing and Answering Binary Visual Questions
Peng Zhang, Yash Goyal, Douglas Summers-Stay +2
cs.CLcs.CVcs.LGarXiv:1511.05099v52015EvalCrafter: Benchmarking and Evaluating Large Video Generation Models
Yaofang Liu, Xiaodong Cun, Xuebo Liu +7
cs.CVarXiv:2310.11440v32023OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents
Hugo Laurençon, Lucile Saulnier, Léo Tronchon +9
cs.IRcs.CVarXiv:2306.16527v22023Earthformer: Exploring Space-Time Transformers for Earth System Forecasting
Zhihan Gao, Xingjian Shi, Hao Wang +4
cs.LGcs.AIcs.CVarXiv:2207.05833v22022ControlVideo: Training-free Controllable Text-to-Video Generation
Yabo Zhang, Yuxiang Wei, Dongsheng Jiang +3
cs.CVarXiv:2305.13077v12023SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery
Xin Guo, Jiangwei Lao, Bo Dang +13
cs.CVarXiv:2312.10115v22023Smart Mining for Deep Metric Learning
Ben Harwood, Vijay Kumar B G, Gustavo Carneiro +2
cs.CVarXiv:1704.01285v32017Video-P2P: Video Editing with Cross-attention Control
Shaoteng Liu, Yuechen Zhang, Wenbo Li +2
cs.CVarXiv:2303.04761v12023SGCN:Sparse Graph Convolution Network for Pedestrian Trajectory Prediction
Liushuai Shi, Le Wang, Chengjiang Long +4
cs.CVarXiv:2104.01528v12021Learning to diagnose from scratch by exploiting dependencies among labels
Li Yao, Eric Poblenz, Dmitry Dagunts +3
cs.CVarXiv:1710.10501v22017Keeping Your Eye on the Ball: Trajectory Attention in Video Transformers
Mandela Patrick, Dylan Campbell, Yuki M. Asano +5
cs.CVarXiv:2106.05392v22021What does a platypus look like? Generating customized prompts for zero-shot image classification
Sarah Pratt, Ian Covert, Rosanne Liu +1
cs.CVcs.LGarXiv:2209.03320v32022Adversarial Spatio-Temporal Learning for Video Deblurring
Kaihao Zhang, Wenhan Luo, Yiran Zhong +3
cs.CVarXiv:1804.00533v22018Sharp U-Net: Depthwise Convolutional Network for Biomedical Image Segmentation
Hasib Zunair, A. Ben Hamza
eess.IVcs.CVarXiv:2107.12461v12021Progressive Domain Adaptation for Object Detection
Han-Kai Hsu, Chun-Han Yao, Yi-Hsuan Tsai +4
cs.CVarXiv:1910.11319v12019VectorMapNet: End-to-end Vectorized HD Map Learning
Yicheng Liu, Tianyuan Yuan, Yue Wang +2
cs.CVcs.ROarXiv:2206.08920v62022Modeling the Background for Incremental Learning in Semantic Segmentation
Fabio Cermelli, Massimiliano Mancini, Samuel Rota Bulò +2
cs.CVarXiv:2002.00718v22020Flexible Diffusion Modeling of Long Videos
William Harvey, Saeid Naderiparizi, Vaden Masrani +2
cs.CVcs.LGarXiv:2205.11495v32022In-Place Activated BatchNorm for Memory-Optimized Training of DNNs
Samuel Rota Bulò, Lorenzo Porzi, Peter Kontschieder
cs.CVarXiv:1712.02616v32017Appearance-and-Relation Networks for Video Classification
Limin Wang, Wei Li, Wen Li +1
cs.CVarXiv:1711.09125v22017UniFormer: Unified Transformer for Efficient Spatiotemporal Representation Learning
Kunchang Li, Yali Wang, Peng Gao +4
cs.CVarXiv:2201.04676v32022Skeleton-Based Action Recognition with Spatial Reasoning and Temporal Stack Learning
Chenyang Si, Ya Jing, Wei Wang +2
cs.CVarXiv:1805.02335v22018Morphing and Sampling Network for Dense Point Cloud Completion
Minghua Liu, Lu Sheng, Sheng Yang +2
cs.CVarXiv:1912.00280v12019Building a Large Scale Dataset for Image Emotion Recognition: The Fine Print and The Benchmark
Quanzeng You, Jiebo Luo, Hailin Jin +1
cs.AIcs.CVarXiv:1605.02677v12016Residual Networks of Residual Networks: Multilevel Residual Networks
Ke Zhang, Miao Sun, Tony X. Han +3
cs.CVarXiv:1608.02908v22016Video Summarization with Attention-Based Encoder-Decoder Networks
Zhong Ji, Kailin Xiong, Yanwei Pang +1
cs.CVarXiv:1708.09545v22017Learning Two-View Correspondences and Geometry Using Order-Aware Network
Jiahui Zhang, Dawei Sun, Zixin Luo +6
cs.CVcs.CGcs.LGarXiv:1908.04964v12019Deep Learning Based Brain Tumor Segmentation: A Survey
Zhihua Liu, Lei Tong, Zheheng Jiang +6
eess.IVcs.CVarXiv:2007.09479v32020Self-supervised Learning in Remote Sensing: A Review
Yi Wang, Conrad M Albrecht, Nassim Ait Ali Braham +2
cs.CVarXiv:2206.13188v22022NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario
Tianwen Qian, Jingjing Chen, Linhai Zhuo +2
cs.CVarXiv:2305.14836v22023Neural Nearest Neighbors Networks
Tobias Plötz, Stefan Roth
cs.CVcs.LGarXiv:1810.12575v12018DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data
Stephanie Fu, Netanel Tamir, Shobhita Sundaram +4
cs.CVcs.LGarXiv:2306.09344v32023Camera Distance-aware Top-down Approach for 3D Multi-person Pose Estimation from a Single RGB Image
Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee
cs.CVarXiv:1907.11346v22019RoadTracer: Automatic Extraction of Road Networks from Aerial Images
Favyen Bastani, Songtao He, Sofiane Abbar +5
cs.CVarXiv:1802.03680v22018Classification of Hyperspectral and LiDAR Data Using Coupled CNNs
Renlong Hang, Zhu Li, Pedram Ghamisi +3
cs.CVeess.IVarXiv:2002.01144v12020Low-rank Bilinear Pooling for Fine-Grained Classification
Shu Kong, Charless Fowlkes
cs.CVarXiv:1611.05109v22016Compressed Video Action Recognition
Chao-Yuan Wu, Manzil Zaheer, Hexiang Hu +3
cs.CVarXiv:1712.00636v22017Neural Prototype Trees for Interpretable Fine-grained Image Recognition
Meike Nauta, Ron van Bree, Christin Seifert
cs.CVcs.AIcs.LGarXiv:2012.02046v22020Uncertainty-Aware Blind Image Quality Assessment in the Laboratory and Wild
Weixia Zhang, Kede Ma, Guangtao Zhai +1
cs.CVcs.LGcs.MMarXiv:2005.13983v62020Disentangled Non-Local Neural Networks
Minghao Yin, Zhuliang Yao, Yue Cao +4
cs.CVcs.CLcs.LGarXiv:2006.06668v22020How Neural Networks Extrapolate: From Feedforward to Graph Neural Networks
Keyulu Xu, Mozhi Zhang, Jingling Li +3
cs.LGcs.AIcs.CVarXiv:2009.11848v52020Stereo DSO: Large-Scale Direct Sparse Visual Odometry with Stereo Cameras
Rui Wang, Martin Schwörer, Daniel Cremers
cs.CVarXiv:1708.07878v12017Designing Deep Networks for Surface Normal Estimation
Xiaolong Wang, David F. Fouhey, Abhinav Gupta
cs.CVarXiv:1411.4958v12014Multi-task Collaborative Network for Joint Referring Expression Comprehension and Segmentation
Gen Luo, Yiyi Zhou, Xiaoshuai Sun +4
cs.CVarXiv:2003.08813v12020Appearance-Based Loop Closure Detection for Online Large-Scale and Long-Term Operation
Mathieu Labbé, François Michaud
cs.ROcs.CVarXiv:2407.15304v12024Dual Motion GAN for Future-Flow Embedded Video Prediction
Xiaodan Liang, Lisa Lee, Wei Dai +1
cs.CVarXiv:1708.00284v22017TOPIQ: A Top-down Approach from Semantics to Distortions for Image Quality Assessment
Chaofeng Chen, Jiadi Mo, Jingwen Hou +5
cs.CVarXiv:2308.03060v12023Plug-and-Play Priors for Bright Field Electron Tomography and Sparse Interpolation
Suhas Sreehari, S. V. Venkatakrishnan, Brendt Wohlberg +3
cs.CVeess.IVarXiv:1512.07331v12015Probabilistic Face Embeddings
Yichun Shi, Anil K. Jain
cs.CVarXiv:1904.09658v42019FaceScape: a Large-scale High Quality 3D Face Dataset and Detailed Riggable 3D Face Prediction
Haotian Yang, Hao Zhu, Yanru Wang +4
cs.CVarXiv:2003.13989v32020Hough-CNN: Deep Learning for Segmentation of Deep Brain Regions in MRI and Ultrasound
Fausto Milletari, Seyed-Ahmad Ahmadi, Christine Kroll +8
cs.CVarXiv:1601.07014v32016PolyGen: An Autoregressive Generative Model of 3D Meshes
Charlie Nash, Yaroslav Ganin, S. M. Ali Eslami +1
cs.GRcs.CVcs.LGarXiv:2002.10880v12020HP-GAN: Probabilistic 3D human motion prediction via GAN
Emad Barsoum, John Kender, Zicheng Liu
cs.CVcs.AIcs.HCarXiv:1711.09561v12017Variational Denoising Network: Toward Blind Noise Modeling and Removal
Zongsheng Yue, Hongwei Yong, Qian Zhao +2
cs.CVarXiv:1908.11314v52019Can Deep Learning Outperform Modern Commercial CT Image Reconstruction Methods?
Hongming Shan, Atul Padole, Fatemeh Homayounieh +5
cs.CVphysics.med-pharXiv:1811.03691v12018UC-Net: Uncertainty Inspired RGB-D Saliency Detection via Conditional Variational Autoencoders
Jing Zhang, Deng-Ping Fan, Yuchao Dai +4
cs.CVarXiv:2004.05763v12020Synergistic Image and Feature Adaptation: Towards Cross-Modality Domain Adaptation for Medical Image Segmentation
Cheng Chen, Qi Dou, Hao Chen +2
cs.CVarXiv:1901.08211v42019DecideNet: Counting Varying Density Crowds Through Attention Guided Detection and Density Estimation
Jiang Liu, Chenqiang Gao, Deyu Meng +1
cs.CVarXiv:1712.06679v22017YOLOv6 v3.0: A Full-Scale Reloading
Chuyi Li, Lulu Li, Yifei Geng +6
cs.CVarXiv:2301.05586v12023