Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
9,121 to 9,180 of 18,817
A Single Stream Network for Robust and Real-time RGB-D Salient Object Detection
Xiaoqi Zhao, Lihe Zhang, Youwei Pang +2
cs.CVarXiv:2007.06811v22020Who Said What: Modeling Individual Labelers Improves Classification
Melody Y. Guan, Varun Gulshan, Andrew M. Dai +1
cs.LGcs.CVarXiv:1703.08774v22017Transfer learning for music classification and regression tasks
Keunwoo Choi, György Fazekas, Mark Sandler +1
cs.CVcs.AIcs.MMarXiv:1703.09179v42017Revisiting Perspective Information for Efficient Crowd Counting
Miaojing Shi, Zhaohui Yang, Chao Xu +1
cs.CVarXiv:1807.01989v32018Contextual Diversity for Active Learning
Sharat Agarwal, Himanshu Arora, Saket Anand +1
cs.CVarXiv:2008.05723v12020Dual Cross-Attention Learning for Fine-Grained Visual Categorization and Object Re-Identification
Haowei Zhu, Wenjing Ke, Dong Li +3
cs.CVcs.AIcs.LGarXiv:2205.02151v12022Vision Permutator: A Permutable MLP-Like Architecture for Visual Recognition
Qibin Hou, Zihang Jiang, Li Yuan +3
cs.CVarXiv:2106.12368v12021Resolution Adaptive Networks for Efficient Inference
Le Yang, Yizeng Han, Xi Chen +3
cs.CVarXiv:2003.07326v52020TCTrack: Temporal Contexts for Aerial Tracking
Ziang Cao, Ziyuan Huang, Liang Pan +3
cs.CVarXiv:2203.01885v32022Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions
Jaewoo Ahn, Junseo Kim, Hyunseo Kim +4
cs.CLcs.AIcs.CVarXiv:2608.30428v12026Learning to Generalize Unseen Domains via Memory-based Multi-Source Meta-Learning for Person Re-Identification
Yuyang Zhao, Zhun Zhong, Fengxiang Yang +4
cs.CVarXiv:2012.00417v32020SGUIE-Net: Semantic Attention Guided Underwater Image Enhancement with Multi-Scale Perception
Qi Qi, Kunqian Li, Haiyong Zheng +3
eess.IVcs.CVarXiv:2201.02832v12022CoVA-SFT: A Large-Scale Dataset for Chain of Visual Abstractions
Tsung-Han Wu, Heekyung Lee, Anya Ji +4
cs.CLcs.CVarXiv:2608.28958v12026Pose2Seg: Detection Free Human Instance Segmentation
Song-Hai Zhang, Ruilong Li, Xin Dong +6
cs.CVarXiv:1803.10683v32018Weaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching
Rong Shan, Tianyi Xu, Congmin Zheng +9
cs.CVcs.IRarXiv:2608.28695v12026Detecting Sarcasm in Multimodal Social Platforms
Rossano Schifanella, Paloma de Juan, Joel Tetreault +1
cs.CVcs.CLcs.MMarXiv:1608.02289v12016Time Will Tell: New Outlooks and A Baseline for Temporal Multi-View 3D Object Detection
Jinhyung Park, Chenfeng Xu, Shijia Yang +4
cs.CVcs.AIcs.LGarXiv:2210.02443v12022Domain Enhanced Arbitrary Image Style Transfer via Contrastive Learning
Yuxin Zhang, Fan Tang, Weiming Dong +4
cs.CVcs.GRarXiv:2205.09542v22022Understanding and Mitigating Copying in Diffusion Models
Gowthami Somepalli, Vasu Singla, Micah Goldblum +2
cs.LGcs.CRcs.CVarXiv:2305.20086v12023PiCIE: Unsupervised Semantic Segmentation using Invariance and Equivariance in Clustering
Jang Hyun Cho, Utkarsh Mall, Kavita Bala +1
cs.CVarXiv:2103.17070v12021Progressive Transformers for End-to-End Sign Language Production
Ben Saunders, Necati Cihan Camgoz, Richard Bowden
cs.CVcs.CLcs.LGarXiv:2004.14874v22020Will we run out of data? Limits of LLM scaling based on human-generated data
Pablo Villalobos, Anson Ho, Jaime Sevilla +3
cs.LGcs.AIcs.CLarXiv:2211.04325v22022Modality to Modality Translation: An Adversarial Representation Learning and Graph Fusion Network for Multimodal Fusion
Sijie Mai, Haifeng Hu, Songlong Xing
cs.CVcs.LGcs.MMarXiv:1911.07848v42019Wave-ViT: Unifying Wavelet and Transformers for Visual Representation Learning
Ting Yao, Yingwei Pan, Yehao Li +2
cs.CVcs.LGarXiv:2207.04978v12022Learning Human Motion Models for Long-term Predictions
Partha Ghosh, Jie Song, Emre Aksan +1
cs.CVarXiv:1704.02827v22017Text-To-4D Dynamic Scene Generation
Uriel Singer, Shelly Sheynin, Adam Polyak +8
cs.CVcs.AIcs.LGarXiv:2301.11280v12023Knowledge Adaptation for Efficient Semantic Segmentation
Tong He, Chunhua Shen, Zhi Tian +3
cs.CVarXiv:1903.04688v12019DeepTAM: Deep Tracking and Mapping
Huizhong Zhou, Benjamin Ummenhofer, Thomas Brox
cs.CVarXiv:1808.01900v22018The Feeling of Success: Does Touch Sensing Help Predict Grasp Outcomes?
Roberto Calandra, Andrew Owens, Manu Upadhyaya +4
cs.ROcs.CVcs.LGarXiv:1710.05512v22017Global canopy height regression and uncertainty estimation from GEDI LIDAR waveforms with deep ensembles
Nico Lang, Nikolai Kalischek, John Armston +3
cs.LGcs.CVarXiv:2103.03975v22021Generative replay with feedback connections as a general strategy for continual learning
Gido M. van de Ven, Andreas S. Tolias
cs.LGcs.AIcs.CVarXiv:1809.10635v22018Towards Effective Low-bitwidth Convolutional Neural Networks
Bohan Zhuang, Chunhua Shen, Mingkui Tan +2
cs.CVarXiv:1711.00205v22017ReVA: A Region-Aware Visual Assistant for Visually Grounded Question Answering
Anoop Senthil
cs.CLcs.CVarXiv:2608.28707v12026RMT: Retentive Networks Meet Vision Transformers
Qihang Fan, Huaibo Huang, Mingrui Chen +2
cs.CVarXiv:2309.11523v62023On learning to localize objects with minimal supervision
Hyun Oh Song, Ross Girshick, Stefanie Jegelka +3
cs.CVcs.LGarXiv:1403.1024v42014Prompt, Generate, then Cache: Cascade of Foundation Models makes Strong Few-shot Learners
Renrui Zhang, Xiangfei Hu, Bohao Li +5
cs.CVcs.CLarXiv:2303.02151v12023Person Re-Identification via Recurrent Feature Aggregation
Yichao Yan, Bingbing Ni, Zhichao Song +3
cs.CVarXiv:1701.06351v12017Learning a Text-Video Embedding from Incomplete and Heterogeneous Data
Antoine Miech, Ivan Laptev, Josef Sivic
cs.CVarXiv:1804.02516v22018SplattingAvatar: Realistic Real-Time Human Avatars with Mesh-Embedded Gaussian Splatting
Zhijing Shao, Zhaolong Wang, Zhuang Li +5
cs.GRcs.CVarXiv:2403.05087v12024Embedding Structured Contour and Location Prior in Siamesed Fully Convolutional Networks for Road Detection
Qi Wang, Junyu Gao, Yuan Yuan
cs.CVarXiv:1905.01575v12019DC-UNet: Rethinking the U-Net Architecture with Dual Channel Efficient CNN for Medical Images Segmentation
Ange Lou, Shuyue Guan, Murray Loew
eess.IVcs.CVarXiv:2006.00414v12020Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory
Runjia Qian, Zile Wang, Jihai Zhang +14
cs.CVarXiv:2608.29910v12026RGBT Salient Object Detection: A Large-scale Dataset and Benchmark
Zhengzheng Tu, Yan Ma, Zhun Li +3
cs.CVarXiv:2007.03262v62020Object Motion Guided Human Motion Synthesis
Jiaman Li, Jiajun Wu, C. Karen Liu
cs.CVarXiv:2309.16237v12023OmniDepth: Dense Depth Estimation for Indoors Spherical Panoramas
Nikolaos Zioulis, Antonis Karakottas, Dimitrios Zarpalas +1
cs.CVarXiv:1807.09620v12018Forgery-aware Adaptive Transformer for Generalizable Synthetic Image Detection
Huan Liu, Zichang Tan, Chuangchuang Tan +3
cs.CVarXiv:2312.16649v12023Abdominal multi-organ segmentation with organ-attention networks and statistical fusion
Yan Wang, Yuyin Zhou, Wei Shen +3
cs.CVarXiv:1804.08414v12018Explainable Medical Imaging AI Needs Human-Centered Design: Guidelines and Evidence from a Systematic Review
Haomin Chen, Catalina Gomez, Chien-Ming Huang +1
cs.HCcs.CVcs.LGarXiv:2112.12596v42021Practical Blind Image Denoising via Swin-Conv-UNet and Data Synthesis
Kai Zhang, Yawei Li, Jingyun Liang +6
cs.CVcs.GReess.IVarXiv:2203.13278v42022Parametric Multimodal User Memory: Storing What Captions Cannot Carry
Bojie Li, Noah Shi
cs.CLcs.AIcs.CVarXiv:2608.28609v12026Batch-Instance Normalization for Adaptively Style-Invariant Neural Networks
Hyeonseob Nam, Hyo-Eun Kim
cs.CVarXiv:1805.07925v32018DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution
Jiashu Zhu, Yanhao Zheng, Ruitian Tian +7
cs.CVcs.SDarXiv:2608.31106v12026Fast Inference in Sparse Coding Algorithms with Applications to Object Recognition
Koray Kavukcuoglu, Marc'Aurelio Ranzato, Yann LeCun
cs.CVcs.LGarXiv:1010.3467v12010Defense Against Adversarial Attacks Using Feature Scattering-based Adversarial Training
Haichao Zhang, Jianyu Wang
cs.CVcs.CRcs.LGarXiv:1907.10764v42019Learning Task-Oriented Grasping for Tool Manipulation from Simulated Self-Supervision
Kuan Fang, Yuke Zhu, Animesh Garg +4
cs.ROcs.CVcs.LGarXiv:1806.09266v12018Image Deformation Meta-Networks for One-Shot Learning
Zitian Chen, Yanwei Fu, Yu-Xiong Wang +3
cs.CVarXiv:1905.11641v22019Shield: Fast, Practical Defense and Vaccination for Deep Learning using JPEG Compression
Nilaksh Das, Madhuri Shanbhogue, Shang-Tse Chen +5
cs.CVcs.AIcs.CRarXiv:1802.06816v12018Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained Models
Guillermo Ortiz-Jimenez, Alessandro Favero, Pascal Frossard
cs.LGcs.CVarXiv:2305.12827v32023Unsupervised Domain Adaptation through Self-Supervision
Yu Sun, Eric Tzeng, Trevor Darrell +1
cs.LGcs.CVstat.MLarXiv:1909.11825v22019HyperSeg: Patch-wise Hypernetwork for Real-time Semantic Segmentation
Yuval Nirkin, Lior Wolf, Tal Hassner
cs.CVarXiv:2012.11582v22020