Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
13,441 to 13,500 of 18,841
Learning Blind Video Temporal Consistency
Wei-Sheng Lai, Jia-Bin Huang, Oliver Wang +3
cs.CVarXiv:1808.00449v12018Neural Point-Based Graphics
Kara-Ali Aliev, Artem Sevastopolsky, Maria Kolos +2
cs.CVarXiv:1906.08240v32019Deformable GANs for Pose-based Human Image Generation
Aliaksandr Siarohin, Enver Sangineto, Stephane Lathuiliere +1
cs.CVarXiv:1801.00055v22017Generating Images from Captions with Attention
Elman Mansimov, Emilio Parisotto, Jimmy Lei Ba +1
cs.LGcs.CVarXiv:1511.02793v22015MPDIoU: A Loss for Efficient and Accurate Bounding Box Regression
Siliang Ma, Yong Xu
cs.CVcs.AIarXiv:2307.07662v12023Recognizing Image Style
Sergey Karayev, Matthew Trentacoste, Helen Han +4
cs.CVarXiv:1311.3715v32013Infrared and Visible Image Fusion using a Deep Learning Framework
Hui Li, Xiao-Jun Wu, Josef Kittler
cs.CVarXiv:1804.06992v42018Automatic Facial Expression Recognition Using Features of Salient Facial Patches
S L Happy, Aurobinda Routray
cs.CVarXiv:1505.04026v12015Grid R-CNN
Xin Lu, Buyu Li, Yuxin Yue +2
cs.CVarXiv:1811.12030v12018Indoor Semantic Segmentation using depth information
Camille Couprie, Clément Farabet, Laurent Najman +1
cs.CVarXiv:1301.3572v22013Deep Label Distribution Learning with Label Ambiguity
Bin-Bin Gao, Chao Xing, Chen-Wei Xie +2
cs.CVarXiv:1611.01731v22016DPOD: 6D Pose Object Detector and Refiner
Sergey Zakharov, Ivan Shugurov, Slobodan Ilic
cs.CVcs.ROarXiv:1902.11020v32019Episodic Training for Domain Generalization
Da Li, Jianshu Zhang, Yongxin Yang +3
cs.CVarXiv:1902.00113v32019Hierarchical Multi-Scale Attention for Semantic Segmentation
Andrew Tao, Karan Sapra, Bryan Catanzaro
cs.CVarXiv:2005.10821v12020ZeroQ: A Novel Zero Shot Quantization Framework
Yaohui Cai, Zhewei Yao, Zhen Dong +3
cs.CVarXiv:2001.00281v12020StyleBank: An Explicit Representation for Neural Image Style Transfer
Dongdong Chen, Lu Yuan, Jing Liao +2
cs.CVarXiv:1703.09210v22017Video Transformer Network
Daniel Neimark, Omri Bar, Maya Zohar +1
cs.CVarXiv:2102.00719v32021PlaNet - Photo Geolocation with Convolutional Neural Networks
Tobias Weyand, Ilya Kostrikov, James Philbin
cs.CVarXiv:1602.05314v12016Neural 3D Video Synthesis from Multi-view Video
Tianye Li, Mira Slavcheva, Michael Zollhoefer +8
cs.CVcs.GRarXiv:2103.02597v22021Gaussian Grouping: Segment and Edit Anything in 3D Scenes
Mingqiao Ye, Martin Danelljan, Fisher Yu +1
cs.CVcs.AIarXiv:2312.00732v22023Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence Maps
Yue Hu, Shaoheng Fang, Zixing Lei +2
cs.CVarXiv:2209.12836v12022Arbitrary Style Transfer with Style-Attentional Networks
Dae Young Park, Kwang Hee Lee
cs.CVarXiv:1812.02342v52018Gaze360: Physically Unconstrained Gaze Estimation in the Wild
Petr Kellnhofer, Adria Recasens, Simon Stent +2
cs.CVarXiv:1910.10088v12019Composing Text and Image for Image Retrieval - An Empirical Odyssey
Nam Vo, Lu Jiang, Chen Sun +4
cs.CVarXiv:1812.07119v12018Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions?
Abhishek Das, Harsh Agrawal, C. Lawrence Zitnick +2
stat.MLcs.CVarXiv:1606.05589v12016CREST: Convolutional Residual Learning for Visual Tracking
Yibing Song, Chao Ma, Lijun Gong +3
cs.CVcs.AIcs.MMarXiv:1708.00225v12017Frustum ConvNet: Sliding Frustums to Aggregate Local Point-Wise Features for Amodal 3D Object Detection
Zhixin Wang, Kui Jia
cs.CVarXiv:1903.01864v22019SEAN: Image Synthesis with Semantic Region-Adaptive Normalization
Peihao Zhu, Rameen Abdal, Yipeng Qin +1
cs.CVcs.GReess.IVarXiv:1911.12861v22019Automatic Ship Detection of Remote Sensing Images from Google Earth in Complex Scenes Based on Multi-Scale Rotation Dense Feature Pyramid Networks
Xue Yang, Hao Sun, Kun Fu +4
cs.CVarXiv:1806.04331v12018AugFPN: Improving Multi-scale Feature Learning for Object Detection
Chaoxu Guo, Bin Fan, Qian Zhang +2
cs.CVarXiv:1912.05384v12019DiffuserCam: Lensless Single-exposure 3D Imaging
Nick Antipa, Grace Kuo, Reinhard Heckel +4
cs.CVarXiv:1710.02134v12017Skeleton-Based Action Recognition Using Spatio-Temporal LSTM Network with Trust Gates
Jun Liu, Amir Shahroudy, Dong Xu +2
cs.CVarXiv:1706.08276v12017Shift-Net: Image Inpainting via Deep Feature Rearrangement
Zhaoyi Yan, Xiaoming Li, Mu Li +2
cs.CVarXiv:1801.09392v22018An Energy and GPU-Computation Efficient Backbone Network for Real-Time Object Detection
Youngwan Lee, Joong-won Hwang, Sangrok Lee +2
cs.CVarXiv:1904.09730v12019The inD Dataset: A Drone Dataset of Naturalistic Road User Trajectories at German Intersections
Julian Bock, Robert Krajewski, Tobias Moers +3
cs.CVcs.LGcs.ROarXiv:1911.07602v12019Noiseprint: a CNN-based camera model fingerprint
Davide Cozzolino, Luisa Verdoliva
cs.CVarXiv:1808.08396v12018TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
Shuhuai Ren, Linli Yao, Shicheng Li +2
cs.CVcs.AIcs.CLarXiv:2312.02051v22023Visual Saliency Transformer
Nian Liu, Ni Zhang, Kaiyuan Wan +2
cs.CVarXiv:2104.12099v22021Rethinking the Value of Labels for Improving Class-Imbalanced Learning
Yuzhe Yang, Zhi Xu
cs.LGcs.CVstat.MLarXiv:2006.07529v22020AppAgent: Multimodal Agents as Smartphone Users
Chi Zhang, Zhao Yang, Jiaxuan Liu +6
cs.CVarXiv:2312.13771v32023Understanding Deep Networks via Extremal Perturbations and Smooth Masks
Ruth Fong, Mandela Patrick, Andrea Vedaldi
cs.CVcs.LGstat.MLarXiv:1910.08485v12019DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models
Zijie J. Wang, Evan Montoya, David Munechika +3
cs.CVcs.AIcs.HCarXiv:2210.14896v42022Shap-E: Generating Conditional 3D Implicit Functions
Heewoo Jun, Alex Nichol
cs.CVcs.LGarXiv:2305.02463v12023SimSwap: An Efficient Framework For High Fidelity Face Swapping
Renwang Chen, Xuanhong Chen, Bingbing Ni +1
cs.CVarXiv:2106.06340v12021Uncertainty Sets for Image Classifiers using Conformal Prediction
Anastasios Angelopoulos, Stephen Bates, Jitendra Malik +1
cs.CVmath.STstat.MLarXiv:2009.14193v52020Graph-Based Global Reasoning Networks
Yunpeng Chen, Marcus Rohrbach, Zhicheng Yan +3
cs.CVarXiv:1811.12814v12018Towards Fast, Accurate and Stable 3D Dense Face Alignment
Jianzhu Guo, Xiangyu Zhu, Yang Yang +3
cs.CVarXiv:2009.09960v22020Sparse Modeling for Image and Vision Processing
Julien Mairal, Francis Bach, Jean Ponce
cs.CVarXiv:1411.3230v22014TransGAN: Two Pure Transformers Can Make One Strong GAN, and That Can Scale Up
Yifan Jiang, Shiyu Chang, Zhangyang Wang
cs.CVarXiv:2102.07074v42021TokenFlow: Consistent Diffusion Features for Consistent Video Editing
Michal Geyer, Omer Bar-Tal, Shai Bagon +1
cs.CVarXiv:2307.10373v32023DF-Net: Unsupervised Joint Learning of Depth and Flow using Cross-Task Consistency
Yuliang Zou, Zelun Luo, Jia-Bin Huang
cs.CVarXiv:1809.01649v12018Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Bowen Shi, Wei-Ning Hsu, Kushal Lakhotia +1
eess.AScs.CVcs.SDarXiv:2201.02184v22022Spatial-Temporal Recurrent Neural Network for Emotion Recognition
Tong Zhang, Wenming Zheng, Zhen Cui +2
cs.CVarXiv:1705.04515v12017Scene Text Detection and Recognition: The Deep Learning Era
Shangbang Long, Xin He, Cong Yao
cs.CVarXiv:1811.04256v52018TransVG: End-to-End Visual Grounding with Transformers
Jiajun Deng, Zhengyuan Yang, Tianlang Chen +2
cs.CVarXiv:2104.08541v42021Illumination-aware Faster R-CNN for Robust Multispectral Pedestrian Detection
Chengyang Li, Dan Song, Ruofeng Tong +1
cs.CVarXiv:1803.05347v22018UniPC: A Unified Predictor-Corrector Framework for Fast Sampling of Diffusion Models
Wenliang Zhao, Lujia Bai, Yongming Rao +2
cs.LGcs.CVarXiv:2302.04867v42023Regressing Robust and Discriminative 3D Morphable Models with a very Deep Neural Network
Anh Tuan Tran, Tal Hassner, Iacopo Masi +1
cs.CVarXiv:1612.04904v12016VL-Adapter: Parameter-Efficient Transfer Learning for Vision-and-Language Tasks
Yi-Lin Sung, Jaemin Cho, Mohit Bansal
cs.CVcs.AIcs.CLarXiv:2112.06825v22021Pluralistic Image Completion
Chuanxia Zheng, Tat-Jen Cham, Jianfei Cai
cs.CVarXiv:1903.04227v22019