Source-linked AI summary

Emergence of grid-like representations by training recurrent neural networks to perform spatial localization

Christopher J. Cueva, Xue-Xin Wei

arXiv:1803.07770v1q-bio.NCcs.AIcs.NEstat.ML

TL;DR

The paper addresses how spatial representations in the entorhinal cortex might arise and trains recurrent neural networks to perform path integration from velocity inputs. The trained networks develop multiple spatial and velocity tuning profiles resembling experimental observations, supporting an efficient neural code for self-location while leaving biological plausibility and spatial-scale diversity unresolved.

  • Problem

    The mechanisms and functional significance of diverse entorhinal spatial representations, including grid cells, remain largely mysterious.

  • Method

    The study trains recurrent neural networks to perform path integration in two-dimensional arenas using the animal’s speed and direction as inputs.

  • Results

    Trained RNNs develop spatial and velocity tuning profiles resembling entorhinal neurophysiology, including grid-like responses and speed tuning.

  • Takeaways & Limitations

    The agreement between model responses and neurophysiology supports the hypothesis that entorhinal populations efficiently represent self-location from velocity input.

  • Takeaways & Limitations

    The learning rule is biologically implausible, and simulations do not reproduce the variety of experimentally observed grid-cell spatial scales.

Abstract

from arXiv · show

Decades of research on the neural code underlying spatial navigation have revealed a diverse set of neural response properties. The Entorhinal Cortex (EC) of the mammalian brain contains a rich set of spatial correlates, including grid cells which encode space using tessellating patterns. However, the mechanisms and functional significance of these spatial representations remain largely mysterious. As a new way to understand these neural representations, we trained recurrent neural networks (RNNs) to perform navigation tasks in 2D arenas based on velocity inputs. Surprisingly, we find that grid-like spatial response patterns emerge in trained networks, along with units that exhibit other spatial correlates, including border cells and band-like cells. All these different functional types of neurons have been observed experimentally. The order of the emergence of grid-like and border cells is also consistent with observations from developmental studies. Together, our results suggest that grid cells, border cells and others as observed in EC may be a natural solution for representing space efficiently given the predominant recurrent connections in the neural circuits.

1 INTRODUCTION

The paper asks how recurrent neural systems might represent self-location during navigation and proposes training RNNs as an alternative to hand-designed spatial models. Trained networks develop grid-like and other spatial response profiles observed in the entorhinal cortex.

  • Spatial navigation requires maintaining self-location and updating it from movement and environmental landmarks.
  • Existing continuous attractor models offer possible mechanisms for grid and place cells but typically rely on finely tuned, systematic connectivity patterns.
  • The study trains an RNN on spatial navigation tasks using biologically relevant constraints and reports grid-like, border-like, and other entorhinal-like responses.
  • The proposed neural representations may provide an efficient way for the brain to represent location during navigation.

2 MODEL

The model is a continuous-time recurrent network that receives speed and direction inputs and learns to estimate two-dimensional position. Training combines localization error with regularization and metabolic-cost penalties across multiple arena geometries.

  • 2.1 MODEL DESCRIPTION: The network contains N = 100 recurrently connected units whose dynamics follow a continuous-time RNN formulation.
  • 2.1 MODEL DESCRIPTION: Two linear readout neurons combine unit firing rates to estimate the animal’s current two-dimensional location.
  • 2.2 INPUT TO THE NETWORK: The RNN receives the animal’s speed and direction at each timestep and outputs integrated x- and y-coordinates.
  • 2.2 INPUT TO THE NETWORK: Boundary handling resamples angular inputs that would move the simulated animal outside the arena, while simulations begin at the arena center.
  • 2.3 TRAINING: Training minimizes squared position error together with penalties on input and output weights, firing rates, and large network parameters.
  • 2.3 TRAINING: The simulations use square, triangular, and hexagonal arenas to assess navigation across different boundary shapes.

3 RESULTS

The trained RNN accurately localizes position and develops diverse spatial and input-tuning profiles resembling several entorhinal response types. Training dynamics, regularization, and boundary interactions shape these representations and support stable long-term error correction.

  • The trained RNN accurately tracks animal paths across square, triangular, and hexagonal arenas.
  • Spatial tuning: Units develop grid-like, border, band-like, and stable non-regular spatial responses, alongside experimentally observed entorhinal correlates.Grid-like fields form regular lattices influenced by boundary shape; band-like responses are often parallel to boundaries.
  • Speed tuning and head direction tuning: Many neurons show linear speed tuning or no speed selectivity, while border cells tend to have almost zero speed tuning.
  • Speed tuning and head direction tuning: A substantial portion of neurons show direction tuning, but the strongest direction-tuned units generally lack clear spatial firing patterns; grid-like units span weak to strong direction tuning.
  • Development of the tuning properties: Border representations emerge early and persist, whereas grid-like responses typically change substantially before appearing later in training.This developmental order is roughly consistent with rodent studies, where border cells emerge before mature grid cells.
  • The importance of regularization: Grid-like representations require regularization that combines noise-robust information storage with metabolic-cost penalties.Grid-like responses were not observed without metabolic cost, and they emerged when speed input was frequently set to zero while Gaussian noise was added.
  • Error correction around the boundary: Over paths several orders of magnitude longer than training sequences, localization error remains stable and spatial response profiles remain stable.Boundary interactions reduce accumulated squared error, with more frequent interactions producing greater error reduction.

4 DISCUSSION

The discussion argues that recurrent networks trained for path integration can reproduce several entorhinal spatial response properties and support an efficient spatial code hypothesis, while emphasizing important limitations.

  • Main findings: RNNs trained for path integration in 2D arenas developed spatial and velocity tuning profiles resembling entorhinal neurophysiology.The reported similarity also includes the timing of distinct neuron-type emergence during training or development.
  • Interpretation: The agreement between model responses and EC recordings supports the hypothesis that EC populations efficiently represent self-location from velocity inputs.
  • Relation to prior work: The recurrent-network approach addresses the brain’s abundant recurrent connectivity, contrasting with prior emphasis on feedforward and convolutional architectures.
  • Relation to prior work: The work is conceptually related to efficient coding in visual processing, including sparse-coding accounts of V1-like filters, but the authors note important differences.
  • Relation to prior work: Unlike feedforward models that derive grid cells from place-cell inputs, this approach addresses how spatial tuning might arise without presupposing place-cell representations.
  • Limitations: The model qualitatively matches EC responses but uses a biologically implausible learning rule and does not reproduce experimentally observed multiple grid spatial scales.The authors identify biologically plausible learning and hierarchical spatial scales as future directions.

A TRIANGULAR ENVIRONMENT

The figure contrasts spatial responses under noise without metabolic cost against responses without noise with metabolic cost.

  • The left side shows spatial responses from a network trained with noise and no metabolic cost.
  • The right side shows spatial responses from a network trained with no noise and metabolic cost.
  • The displayed activity range is labeled from -1 to 1.

B RECTANGULAR ENVIRONMENT

The supplied passage contains an activity-of-unit label with a range from -1 to 1.

  • The passage labels unit activity as ranging from -1 to 1.
  • The passage provides no additional textual description of the rectangular environment.
  • No comparison or outcome is stated in the supplied passage.

C HEXAGONAL ENVIRONMENT

The supplied passage contains an activity-of-unit label but no accompanying description of the hexagonal environment.

  • The passage labels unit activity without specifying an environmental result.
  • The displayed passage does not state axes, conditions, or comparisons for the hexagonal environment.
  • No outcome is reported in the supplied passage.

D RELATION BETWEEN SPEED, DIRECTION, AND SPATIAL SELECTIVITY

The section defines separate measures for speed, direction, and spatial selectivity of individual units, with zero indicating no modulation by the corresponding variable.

  • Speed selectivity is the absolute slope of a fitted line through each unit’s activity tuning curve as a function of speed.
  • Direction selectivity is the difference between the maximum and minimum of each unit’s average activity across input directions.
  • Spatial selectivity is quantified using lifetime sparseness, with each plotted dot representing one unit’s selectivity.

E ADDITIONAL TRAINING DETAILS

Training balanced the error and regularization objectives by scheduling the L2 regularization weight and adaptively controlling the recurrent regularization weight.

  • Training balanced the error function E, L2 regularization RL2, and recurrent regularization RF R so no term dominated or was neglected.
  • The RL2 weight initially matched E and then decreased according to the schedule used by Martens & Sutskever (2011).
  • The RF R weight started at E/10 and was adaptively adjusted during training with an upper bound of E/3.
Loading 1803.07770v1…