Every paper with a summary

Every arXiv paper Paperlayer has summarized so far. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

57,661 to 57,720 of 61,270

  1. PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume

    Deqing Sun, Xiaodong Yang, Ming-Yu Liu +1

    cs.CVarXiv:1709.02371v32017
  2. Self-Attention with Relative Position Representations

    Peter Shaw, Jakob Uszkoreit, Ashish Vaswani

    cs.CLarXiv:1803.02155v22018
  3. Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

    Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Deepak Gotmare +3

    cs.CVcs.AIarXiv:2107.07651v22021
  4. Prompt-to-Prompt Image Editing with Cross Attention Control

    Amir Hertz, Ron Mokady, Jay Tenenbaum +3

    cs.CVcs.CLcs.GRarXiv:2208.01626v12022
  5. Temporal Ensembling for Semi-Supervised Learning

    Samuli Laine, Timo Aila

    cs.NEcs.LGarXiv:1610.02242v32016
  6. Convolutional Pose Machines

    Shih-En Wei, Varun Ramakrishna, Takeo Kanade +1

    cs.CVarXiv:1602.00134v42016
  7. G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment

    Yang Liu, Dan Iter, Yichong Xu +3

    cs.CLcs.AIarXiv:2303.16634v32023
  8. Unsupervised Learning of Depth and Ego-Motion from Video

    Tinghui Zhou, Matthew Brown, Noah Snavely +1

    cs.CVarXiv:1704.07813v22017
  9. A Survey on Deep Transfer Learning

    Chuanqi Tan, Fuchun Sun, Tao Kong +3

    cs.LGstat.MLarXiv:1808.01974v12018
  10. Image-based Recommendations on Styles and Substitutes

    Julian McAuley, Christopher Targett, Qinfeng Shi +1

    cs.CVcs.IRarXiv:1506.04757v12015
  11. StackGAN: Text to Photo-realistic Image Synthesis with Stacked Generative Adversarial Networks

    Han Zhang, Tao Xu, Hongsheng Li +4

    cs.CVcs.AIstat.MLarXiv:1612.03242v22016
  12. UNETR: Transformers for 3D Medical Image Segmentation

    Ali Hatamizadeh, Yucheng Tang, Vishwesh Nath +5

    eess.IVcs.CVcs.LGarXiv:2103.10504v32021
  13. Swin Transformer V2: Scaling Up Capacity and Resolution

    Ze Liu, Han Hu, Yutong Lin +9

    cs.CVarXiv:2111.09883v22021
  14. CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning

    Justin Johnson, Bharath Hariharan, Laurens van der Maaten +3

    cs.CVcs.CLcs.LGarXiv:1612.06890v12016
  15. Microsoft COCO Captions: Data Collection and Evaluation Server

    Xinlei Chen, Hao Fang, Tsung-Yi Lin +4

    cs.CVcs.CLarXiv:1504.00325v22015
  16. BERTopic: Neural topic modeling with a class-based TF-IDF procedure

    Maarten Grootendorst

    cs.CLarXiv:2203.05794v12022
  17. MatConvNet - Convolutional Neural Networks for MATLAB

    Andrea Vedaldi, Karel Lenc

    cs.CVcs.LGcs.MSarXiv:1412.4564v32014
  18. LSST: from Science Drivers to Reference Design and Anticipated Data Products

    Željko Ivezić, Steven M. Kahn, J. Anthony Tyson +310

    astro-pharXiv:0805.2366v52008
  19. BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

    BigScience Workshop, :, Teven Le Scao +391

    cs.CLarXiv:2211.05100v42022
  20. LXMERT: Learning Cross-Modality Encoder Representations from Transformers

    Hao Tan, Mohit Bansal

    cs.CLcs.CVcs.LGarXiv:1908.07490v32019
  21. Direct Sparse Odometry

    Jakob Engel, Vladlen Koltun, Daniel Cremers

    cs.CVarXiv:1607.02565v22016
  22. T-GCN: A Temporal Graph ConvolutionalNetwork for Traffic Prediction

    Ling Zhao, Yujiao Song, Chao Zhang +5

    cs.LGstat.MLarXiv:1811.05320v32018
  23. Language Modeling with Gated Convolutional Networks

    Yann N. Dauphin, Angela Fan, Michael Auli +1

    cs.CLarXiv:1612.08083v32016
  24. Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization

    Qian Kou, Xiaofeng Shi, Xiaosong Qiu +1

    cs.CLcs.AIarXiv:2608.20281v12026
  25. Deep Learning with Coherent Nanophotonic Circuits

    Yichen Shen, Nicholas C. Harris, Scott Skirlo +8

    physics.opticsphysics.comp-pharXiv:1610.02365v12016
  26. DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

    Hao Zhang, Feng Li, Shilong Liu +5

    cs.CVarXiv:2203.03605v42022
  27. GOAG: Generative and Object-Agnostic Grasp Planner for Dexterous Robotic Manipulation

    Julien Merand, Boris Meden, Mathieu Grossard +1

    cs.ROcs.AIarXiv:2608.19759v12026
  28. Recurrent Neural Network Regularization

    Wojciech Zaremba, Ilya Sutskever, Oriol Vinyals

    cs.NEarXiv:1409.2329v52014
  29. Reformer: The Efficient Transformer

    Nikita Kitaev, Łukasz Kaiser, Anselm Levskaya

    cs.LGcs.CLstat.MLarXiv:2001.04451v22020
  30. CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning

    Julien Merand, Boris Meden, Liming Chen +1

    cs.ROcs.AIarXiv:2608.19776v12026
  31. VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation

    Yinming Huang, Shuyuan Tu, Xi Yan +5

    cs.CVarXiv:2608.18607v22026
  32. CNN Architectures for Large-Scale Audio Classification

    Shawn Hershey, Sourish Chaudhuri, Daniel P. W. Ellis +10

    cs.SDcs.LGstat.MLarXiv:1609.09430v22016
  33. CosFace: Large Margin Cosine Loss for Deep Face Recognition

    Hao Wang, Yitong Wang, Zheng Zhou +5

    cs.CVarXiv:1801.09414v22018
  34. Efficient Neural Architecture Search via Parameter Sharing

    Hieu Pham, Melody Y. Guan, Barret Zoph +2

    cs.LGcs.CLcs.CVarXiv:1802.03268v22018
  35. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

    Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas +1

    cs.LGstat.MLarXiv:2006.16236v32020
  36. OpenVLA: An Open-Source Vision-Language-Action Model

    Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti +15

    cs.ROcs.LGarXiv:2406.09246v32024
  37. Towards Real-Time and Adaptable LiDAR Scene Completion

    Azhar Hussian, Martin Vossiek, Vasileios Belagiannis

    cs.CVarXiv:2608.16490v12026
  38. Progressive Neural Networks

    Andrei A. Rusu, Neil C. Rabinowitz, Guillaume Desjardins +5

    cs.LGarXiv:1606.04671v42016
  39. SPK: Eliciting Structured Prior Knowledge for Interpretable Out-of-Distribution Detection in Real-Time Object Detection

    Changshun Wu, Weicheng He, Xiaowei Huang +1

    cs.CVcs.LGarXiv:2608.19080v12026
  40. Unsupervised Visual Representation Learning by Context Prediction

    Carl Doersch, Abhinav Gupta, Alexei A. Efros

    cs.CVarXiv:1505.05192v32015
  41. Evaluating Music Context Preservation: A Multi-facet Framework for Music Editing Systems

    Yash Vishe, Eric Xue, Xunyi Jiang +4

    cs.SDcs.AIarXiv:2512.14629v22025
  42. Cyclical Learning Rates for Training Neural Networks

    Leslie N. Smith

    cs.CVcs.LGcs.NEarXiv:1506.01186v62015
  43. $τ_0$-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation

    Xiaowei Cai, Yunuo Cai, Bingao Chen +36

    cs.ROarXiv:2608.16885v12026
  44. Improving Neural Machine Translation Models with Monolingual Data

    Rico Sennrich, Barry Haddow, Alexandra Birch

    cs.CLarXiv:1511.06709v42015
  45. PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents

    Seongjae Kang, Taehyung Yu, Sung Ju Hwang

    cs.AIcs.CLcs.LGarXiv:2608.19861v12026
  46. Session-based Recommendations with Recurrent Neural Networks

    Balázs Hidasi, Alexandros Karatzoglou, Linas Baltrunas +1

    cs.LGcs.IRcs.NEarXiv:1511.06939v42015
  47. FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills

    Zeyu Ren, Ling Yue, Ran Li +5

    cs.AIarXiv:2607.21596v22026
  48. PaLM-E: An Embodied Multimodal Language Model

    Danny Driess, Fei Xia, Mehdi S. M. Sajjadi +19

    cs.LGcs.AIcs.ROarXiv:2303.03378v12023
  49. EXIMO: VLM Guided Exploration of VLA Policies

    Bhavya Sukhija, Oliver Groth, Mohit Shridhar +5

    cs.AIarXiv:2608.19891v12026
  50. DOTA: A Large-scale Dataset for Object Detection in Aerial Images

    Gui-Song Xia, Xiang Bai, Jian Ding +6

    cs.CVarXiv:1711.10398v32017
  51. Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners

    Umberto Cappellazzo, Xubo Liu, Stavros Petridis +1

    eess.AScs.AIcs.SDarXiv:2608.19863v12026
  52. Cross-lingual Language Model Pretraining

    Guillaume Lample, Alexis Conneau

    cs.CLarXiv:1901.07291v12019
  53. Big Bird: Transformers for Longer Sequences

    Manzil Zaheer, Guru Guruganesh, Avinava Dubey +8

    cs.LGcs.CLstat.MLarXiv:2007.14062v22020
  54. NTU RGB+D: A Large Scale Dataset for 3D Human Activity Analysis

    Amir Shahroudy, Jun Liu, Tian-Tsong Ng +1

    cs.CVarXiv:1604.02808v12016
  55. FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving

    Qihang Fan, Huaibo Huang, Zhiying Wu +2

    cs.CLarXiv:2608.19758v12026
  56. A Large Dataset to Train Convolutional Networks for Disparity, Optical Flow, and Scene Flow Estimation

    Nikolaus Mayer, Eddy Ilg, Philip Häusser +4

    cs.CVcs.LGstat.MLarXiv:1512.02134v12015
  57. A Survey on Multi-Task Learning

    Yu Zhang, Qiang Yang

    cs.LGcs.AIarXiv:1707.08114v32017
  58. Fine-Grained Visual Classification of Aircraft

    Subhransu Maji, Esa Rahtu, Juho Kannala +2

    cs.CVarXiv:1306.5151v12013
  59. SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?

    Zhipeng Xu, Jiahao Lu, Yining Zheng +2

    cs.CLcs.SEarXiv:2608.19799v12026
  60. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

    Mohammad Shoeybi, Mostofa Patwary, Raul Puri +3

    cs.CLarXiv:1909.08053v42019