Source-linked AI summary
Polyhedral Geometry of Time-to-First-Spike Neural Networks
Manjot Singh, Guido Montúfar, Gitta Kutyniok
TL;DR
Spiking neural network에 대한 이론적 이해는 여전히 충분히 발전하지 않았으며, 특히 time-to-first-spike 연산이 입력을 어떻게 분할하는지는 아직 명확하지 않다. 이 논문은 causal region의 polyhedral theory를 전개하고, TTFS network가 기존 ReLU network보다 더 풍부하고 깊이에 따라 조합 가능한 분할을 형성할 수 있음을 보인다.
문제
Causal constraint가 입력 공간 분할을 어떻게 형성하는지를 포함해, spiking neural network에 대한 이론적 이해는 여전히 충분히 발전하지 않았다.
방법
이 논문은 firing-time map과 causal region을 polyhedral하게 특성화한 뒤, shallow·deep·shared-weight TTFS network에 대한 region 수의 bound를 도출한다.
결과
TTFS neuron은 구조화된 affine piece를 지수적으로 많이 가지며, deep network는 깊이에 따라 지수적인 region growth를 달성하고 shared weight는 더 정밀한 count를 제공한다.
핵심 시사점 및 한계
Causal geometry는 TTFS architecture와 weight sharing이 neural-network expressivity를 어떻게 제한하거나 확장하는지 설명하는 framework를 제공한다.
핵심 시사점 및 한계
Shared-weight network에 대해서는 깊이에 따라 지수적으로 증가하는 lower bound를 확립하지 못했으며, folding construction에도 장애가 있다.
Abstract
from arXiv · showhide
We study the expressivity of spiking neural networks, which provide a natural framework for asynchronous, event-driven computation complementary to conventional feedforward neural networks. We consider the time-to-first-spike model in a setting for which the input-output map is continuous and piecewise linear, with affine pieces governed by causal feasibility constraints that determine which presynaptic spikes occur before a neuron fires. We first show that each neuron's firing time admits a maxout-like representation with exponentially many, highly constrained affine pieces. We then formalize causal regions as polyhedral regions with fixed causal sets and derive upper and lower bounds on the maximal number of causal regions in both shallow and multilayer feedforward spiking networks. Our theoretical and experimental results show that spiking networks can generate richer partitions of the input space than conventional feedforward ReLU networks.
1 서론
이 연구는 positive-weight, continuous piecewise-linear TTFS spiking network에서 causal region의 기하를 다루는 이론을 전개한다. 단일 neuron의 region을 특성화하고 shallow 및 deep network의 bound를 도출하며, architecture와 initialization에 따른 region complexity를 보이는 실험으로 이를 보완한다.
- 동기: TTFS network는 spike timing으로 output을 encode하지만, causal-region geometry는 piecewise-linear ANN의 activation-region theory에 비해 아직 충분히 발전하지 않았다.각 neuron이 최대 하나의 spike를 emit하고 positive weight를 사용하는 feedforward network에 초점을 맞추며, discontinuous negative-weight regime은 향후 연구로 남겨 둔다.
- 기여: 단일 neuron에서 causal region은 causal constraint로 결정되며, lifted TTFS polytope와 regular subdivision을 통해 polyhedral하게 표현된다.관련 hyperplane arrangement는 실현 가능한 causal set에 대응하는 region의 union으로 구성된다.
- 기여: Shallow network에 대해 TTFS hyperplane arrangement의 upper bound와 실현 가능한 causal region의 점근적 upper 및 lower bound를 도출한다.Weight를 공유하면 causal set이 nested가 되므로, exact region count로부터 더 날카로운 lower bound를 얻을 수 있다.
- 기여: Deep network에서는 folding mechanism에 의해 depth에 따라 exponentially many causal region이 생성되는 반면, layer-shared weight는 subsequent layer의 causal set에 prefix constraint를 부과한다.이러한 constraint는 대응하는 upper bound를 뒷받침하고 실현 가능 및 실현 불가능한 causal pattern을 특성화한다.
- 실험: CIFAR-10과 MNIST에 대한 실험은 estimated SNN region complexity가 width, depth, initialization, ReLU matching, shared-weight restriction에 따라 달라짐을 보인다.실험은 randomly initialized network에 초점을 맞추며, SNN이 많은 causal region을 실현할 수 있음을 확인한다.
2 관련 연구
기존 연구는 activation-region count와 polyhedral geometry를 통해 neural-network expressivity를 측정해 왔으며, SNN 연구에서는 approximation과 computational capability에 관한 결과가 확립되어 있다. SNN의 region-complexity analysis는 상대적으로 부족하며, 특히 causal constraint를 반영한 systematic bound는 드물다.
- ANN의 polyhedral geometry: Activation-region count는 piecewise-linear network가 affine function을 구현하는 영역으로 입력을 분할해 expressivity를 정량화하며, architecture와 depth-width를 비교할 수 있게 한다 [13] [14] [25].
- ANN의 polyhedral geometry: Polyhedral method는 hyperplane arrangement와 bent hyperplane을 통해 ReLU boundary를 기술하며, higher-rank Maxout unit은 여러 affine piece를 통해 더 풍부한 subdivision을 생성한다.
- SNN의 expressivity: SNN expressivity research는 coding scheme과 neuron model을 아우르며, rate-based universality, discontinuous piecewise-linear realization, ReLU emulation bound, positive-weight TTFS approximation 및 generalization result를 포함한다 [22] [23].
- SNN의 expressivity: ANN-SNN conversion result는 서로 다른 computational capability를 입증하지 못하므로, spike-time mechanism과 생물학적 동기를 반영한 task에 기반한 비교가 필요하다.
- SNN의 expressivity: ReLU ANN과 비교하면 SNN의 region-complexity analysis는 부족하다. 기존 연구는 상이한 region structure를 보이고 causal set을 연구했지만, systematic shallow- and deep-network counting bound는 아직 제시되지 않았다 [22].
3 TTFS SNN 모델
이 절에서는 single-spike regime의 feedforward TTFS SNN을 정의한다. 여기서 input firing times는 layered spike-time dynamics를 통해 output firing times로 매핑된다. positive weights, nonnegative delays, positive thresholds, ReLU responses를 가정하면 realization은 causal sets가 정하는 affine pieces를 갖는 continuous piecewise linear 함수가 된다.
- 모델 정의: TTFS coding은 실수값 firing times로 정보를 표현하며, 각 neuron은 최대 하나의 spike를 방출하고 layered network는 이러한 spike-time map을 합성한다.realization은 input-neuron firing times를 output-neuron firing times로 매핑한다.
- Spike-time dynamics: Neuron은 누적 potential이 threshold에 처음 도달할 때 fire하며, positive weights는 incoming synapse를 가진 모든 non-input neuron에 대해 unique finite firing time을 보장한다.potential은 continuous하고 nondecreasing하며, 첫 presynaptic spike가 도착한 뒤에는 strictly increasing해진다.
- 모델 정의: 이 모델은 연속적으로 연결된 feedforward layer로 구성되며, layer map을 통해 membrane potential과 firing time을 재귀적으로 정의하고, 이 map들의 composition이 network realization을 이룬다.각 layer map은 이전 layer의 spike-time vector를 다음 layer의 firing-time vector로 변환한다.
- 모델 가정: positive weights, nonnegative delays, positive thresholds, ReLU responses를 사용하면 TTFS realization은 continuous piecewise linear이 된다.ReLU responses는 causal spike-time structure를 보존하면서 input-output map을 piecewise linear하게 만든다.
- Causal regions: ReLU activation pattern과 달리 TTFS affine pieces는 firing 전에 도착하는 presynaptic spike를 기록하는 causal sets로 색인되며, 하나의 causal region은 여러 arrangement cell을 하나로 묶을 수 있다.같은 layer의 서로 다른 neuron은 서로 다른 causal sets를 가질 수도 있다.
4 TTFS 뉴런의 다면체 기하
TTFS 뉴런의 발화 시간 사상은 입력 공간을 다면체 causal region으로 분할하며, 각 region에서 affine하다. 이 지수적으로 풍부한 기하는 제약된 lifted polytope와 dual regular subdivision으로 포착된다.
- Causal region: Positive weight와 threshold는 모든 입력에 대해 unique firing time과 공집합이 아닌 causal set을 보장한다.causal set은 postsynaptic neuron보다 엄 strictly earlier에 발화하는 presynaptic neuron을 포함한다.
- Causal region: 각 causal region은 convex polyhedron이며, firing-time map은 유한 affine function들의 pointwise minimum이므로 continuous, piecewise-affine, and concave하다.이 map은 모든 causal region에서 affine하다.
- Causal region: 단일 neuron은 최대 2^d − 1 causal region을 가지며, positive weight는 모든 공집합이 아닌 causal subset을 feasible하게 만들어 이 bound를 달성한다.따라서 causal region의 수는 input의 수에 따라 지수적으로 증가할 수 있다.
- Arrangement와 causal pattern: TTFS hyperplane arrangement는 causal partition을 세분화하며, translation invariance는 causal geometry가 (d−1)-dimensional quotient에서 relative spike times에만 의존함을 보인다.Multilayer network에서는 arrangement cell이 causal region을 엄 strict하게 과도 세분화할 수 있으므로 causal pattern이 간결한 기술을 제공한다.
- Lifted polytope와 dual complex: Lifted TTFS polytope는 공집합이 아닌 causal subset으로 색인된 2^d − 1 vertices를 가지며, 그 upper hull은 regular subdivision을 통해 causal-region decomposition과 dual이다.Subdivision cell은 I ⊆ J로 색인되는 cube이며 dimension은 |J|−|I|이다. Weight는 boundary orientation을, threshold는 offset을 결정한다.
5 네트워크 수준의 인과 패턴
네트워크 수준의 조합적 대상은 인과 패턴(causal pattern)으로, 층별로 각 뉴런의 presynaptic causal set을 기록한다. 인과 패턴을 고정하면 모든 firing time이 affine인 convex polyhedral input region이 정해진다.
- 네트워크 수준의 인과 패턴: 인과 패턴은 모든 뉴런의 causal set을 층별로 모아, 각 firing time에 영향을 줄 만큼 일찍 도착하는 presynaptic spike가 무엇인지 나타낸다.이는 네트워크의 계산과 정보 전파를 나타내는 조합적 기술자다.
- 네트워크 수준의 인과 패턴: 고정된 각 인과 패턴에 대해 대응하는 input region은 affine consistency constraint로 정의되는 convex polyhedron이다.패턴을 고정하면 모든 뉴런의 causal set이 고정되므로, 네트워크의 재귀 조건을 affine halfspace를 통해 표현할 수 있다.
- 네트워크 수준의 인과 패턴: 각 causal-pattern region 안에서 모든 뉴런의 firing time은 network input의 affine function이다.이 affine expression은 네트워크 층을 거치며 firing time이 전파되는 과정에서 재귀적으로 얻어진다.
6 얕은 SNN의 인과 영역 복잡도
입력 차원이 고정되면 얕은 SNN의 인과 영역 복잡도는 width에 대해 다항식으로 증가하며, 점근적 거동은 Θ(md−1)로 tight하다. hidden neuron들이 weight를 공유하면 nested causal set을 통해 정확한 계수와 더 날카로운 bound를 얻을 수 있으며, 이는 ReLU 영역 scaling과 대비된다.
- 임의의 양의 weight: 단일 neuron은 2d −1개의 공집합이 아닌 causal set을 허용하지만, shared input과 기하학적 제약 때문에 hidden neuron들 사이에서 임의의 조합은 불가능하다.d = 2이고 m = 2일 때 causal-set pair ({1}, {2})와 ({2}, {1})은 불가능하므로, 3^2개의 pair가 모두 나타나지는 않는다.
- 임의의 양의 weight: d가 고정되면 얕은 SNN에서 실현 가능한 causal region의 최대 개수는 neuron당 causal set이 지수적으로 많음에도 Θ(md−1)로 tight하다.upper bound와 constructive lower bound는 hidden-layer width m에 따른 다항식 증가를 확립한다.
- Shared weight: positive weight를 공유하고 threshold가 정렬되어 있으면 causal set은 nested chain을 이루므로, 정확한 열거와 얕은 및 deep network에 대한 개선된 bound가 가능하다.threshold가 증가하면 firing time이 엄격히 증가하므로, θ1 < θ2일 때 S(θ1) ⊆ S(θ2)이다.
- Shared weight: shared-weight shallow network의 최대값은 (m + 1)d − md = (1 − o(1))(m + 1)d as d →∞이다.정확한 개수는 pairwise distinct threshold를 갖는 m개의 hidden neuron과 공통의 positive weight vector에서 달성된다.
- 비교: d가 고정되면 shallow ReLU-ANN 영역은 md만큼 증가하는 반면 SNN causal tuple은 Θ(md−1)만큼 증가하며, m이 고정되면 ReLU saturation 이후 SNN 개수는 d에 대해 exponentially 증가한다.임의의 positive weight는 서로 다른 방향의 causal boundary를 허용하지만, shared weight는 동일한 방향과 threshold로 유도된 nesting을 부과한다.
7 깊은 SNN의 causal region 복잡도
이 절에서는 깊은 SNN에서 causal-pattern 복잡도의 일반적 upper bound, 깊이에 따라 증가하는 constructive lower bound, 그리고 shared weight 조건에서 더 날카로운 bound를 확립한다. 깊이 의존적 구성은 causal region을 공통 output set 위로 반복해서 접으며, shared weight는 이와 같은 exponential growth 메커니즘을 방해한다.
- 일반적 upper bound: 임의의 positive weight에 대해 deep network의 causal pattern 수는 layerwise factor의 유한 곱으로 bound되지만, 이 bound는 느슨할 수 있다.이는 해당 곱이 한 layer 안 뉴런들의 causal set을 독립적으로 변할 수 있다고 취급하기 때문에 발생하며, 실제로는 공통 presynaptic input이 공유 제약을 부과한다.
- Folding 구성: folding 메커니즘은 first layer에서 서로 다른 causal piece를 만들고, second layer에서 firing time을 결합해 반복적인 affine fold를 만든다.각 region에서 relative firing-time map은 동일한 interval로 가는 affine bijection이므로, 이후 block은 region 수를 곱셈적으로 늘릴 수 있다.
- 미해결 확장: parallel fold 또는 두 spike-time ordering을 통한 잠재적 강화는 strictly positive weight나 shared weight 조건에서 여전히 미해결이다.Parallel Cartesian-product fold와 negative relative-time region은 복잡도를 높일 수 있지만, cross-pair coupling, 재사용 가능한 output set, iteration 제약을 해결하려면 새로운 구성이 필요하다.
- Shared weight: shared positive weight와 ordered threshold에서는 각 layer의 output이 임의의 causal tuple이 아니라 lower-dimensional family를 이루므로 recursive combinatorial bound가 더 날카롭다.shared-weight 설정에서는 exponential depth growth가 가능하지 않을 수 있다. 관련된 two-layer 구성은 folding에 필요한 alternating sawtooth 대신 monotone relative output time을 갖기 때문이다.
8 실험
초기화 시 SNN은 CIFAR-10과 MNIST에서 대응하는 ReLU 네트워크보다 trajectory 기반 linear region을 더 많이 보이며, region 수는 width에 따라 빠르게 증가하고 depth와 architecture에 따라 달라진다. 학습 중 정확한 2차원 열거를 수행한 결과, SNN은 대응하는 ReLU 네트워크보다 더 세밀한 local partition을 형성하는 것으로 나타났다.
- Region complexity 추정: 실험에서는 초기화 시 보간된 CIFAR-10 및 MNIST 샘플을 따라 causal-region complexity를 추정하고, 다섯 개 random seed에 대해 결과를 평균냈다.SNN region은 서로 다른 causal pattern으로 식별하며, trajectory에는 균일한 간격의 20,000개 구간을 사용한다.
- Width와 depth의 영향: Region 수는 width에 따라 빠르게 증가하며, 두 데이터셋 모두에서 작은 width에서는 depth의 영향이 작지만 큰 width에서는 depth가 region 수를 늘린다.Depth의 영향은 architecture에 따라 달라지는데, 충분히 넓은 layer가 더 높은 차원의 affine image를 보존하기 때문일 수 있다.
- Independent weight와 shared weight 비교: Shared-weight SNN과 arbitrary-positive-weight SNN은 전반적으로 유사한 region 수를 내지만, 일부 depth와 architecture에서는 arbitrary positive weight가 더 많은 region을 생성한다.이러한 정성적 유사성은 두 데이터셋에서 single-hidden-layer width sweep과 depth 실험 모두에 나타난다.
- SNN과 ReLU 네트워크 비교: SNN은 CIFAR-10과 MNIST의 shallow 및 deep architecture 전반에서 대응하는 ReLU 네트워크보다 일관되게 더 많은 estimated region을 생성한다.이러한 이점은 shallow network에서도 나타나며 depth가 증가해도 뚜렷하게 유지된다.
- 2차원 affine slice에 대한 정확한 열거: 학습 중 정확한 2차원 열거를 수행하면, SNN이 개별 trajectory를 따라 형성하는 것보다 더 풍부한 local partition을 만들고 대응하는 ReLU 네트워크보다 더 세밀한 partition을 형성한다는 사실이 드러난다.분석에는 bounded two-dimensional affine slice를 사용하며, 초기 training epoch에서 depth 3, width 10인 network를 예시로 제시한다.
9 결론 · A causal sets와 causal regions에 관한 보충 세부사항
이 논문은 TTFS 네트워크를 위한 causal-set 및 polyhedral framework를 개발하고, 구조화된 affine regions와 이들이 depth에 따라 어떻게 조합되는지를 규명한다. 결론에서는 causal partitions와 ReLU activation patterns를 대조하고, initialization 실험을 보고하며, delays와 expected region complexity에 관한 미해결 문제를 제시한다.
- 9 결론: 이론적 framework는 고정된 causal sets를 갖는 polyhedral regions를 통해 TTFS input-output maps를 기하학적·조합론적으로 기술한다.이 표현은 asynchronous spike-time dynamics에 맞춰져 있으며, causal feasibility를 중심으로 네트워크의 linear-region structure를 조직한다.
- 9 결론: TTFS firing times는 지수적으로 많은 구조화된 affine functions의 pointwise minima이며, d inputs를 갖는 뉴런에서는 최대 2d −1 causal sets가 가능하다.Causal sets는 firing에 영향을 줄 만큼 일찍 도착하는 presynaptic spikes를 식별하므로, 동일한 input 수를 갖는 ReLU neuron과 구별되는 computational profile을 만든다.
- 9 결론: Deep TTFS layers는 input space를 접고 linear regions를 증식시킬 수 있어, 개별 뉴런의 causal constraints에도 불구하고 depth에 따라 exponential growth를 만든다.Shared weights를 사용하면 뉴런 간 causal sets가 nested가 되어 partition geometry에 추가적인 structural constraint가 생긴다.
- 9 결론: TTFS와 ReLU networks는 모두 CPWL maps를 계산하지만, TTFS regions는 causal patterns를 따르고 ReLU regions는 binary active-or-inactive activation patterns를 따른다.서로 다른 region-generation mechanisms는 근본적으로 다른 combinatorial structures를 만든다.
- 9 결론: Initialization 시 TTFS networks는 training 전에 많은 linear regions를 실현할 수 있어, 대응하는 ReLU architectures와의 비교를 촉진한다.실험은 initialization에서의 region geometry를 조사하고, MNIST와 CIFAR-10에서 width와 depth에 걸쳐 SNNs와 ReLU networks를 비교한다.
- 9 결론: 향후 연구에는 polyhedral representation을 사용해 upper and lower bounds를 더 정밀화하고, TTFS maxout-like structures가 polytope constructions와 어떻게 연관되는지 연구하는 일이 포함된다.제안된 방향은 maxout analyses [29] [31]에서 iterated Minkowski sums와 convex hulls를 통해 형성되는 polytopes의 vertices를 활용한다.
- 9 결론: Delay-dependent causal geometry는 여전히 조건부이다. Theorem 7.3은 nested causal sets를 위해 increasing thresholds를 사용하며, zero-delay setting에서는 공통의 ordered presynaptic firing times를 가정한다.Synaptic delays가 geometry에 미치는 영향은 그 parameterization과 이러한 structural properties가 유지되는지 여부에 크게 좌우된다.
- 9 결론: 이 논문의 bounds는 maximal region complexity에 관한 것으로, 실제로 실현되는 complexity보다 클 수 있다. Random initialization 및 training 후의 expected-region analyses는 여전히 미해결이다.이러한 분석은 causal geometry를 실제 네트워크 동작 및 architectural choices와 더욱 직접적으로 연결할 것이다.
A.1 Lemma 4.1의 증명
각 causal region은 (7)의 affine inequality들이 이루는 고정된 부호 패턴 안에 있으므로 causal set이 일정하게 유지됨을 보인다.
- A.1 Lemma 4.1의 증명: (7)의 각 inequality는 t에 대해 affine이며, 그 경계는 ATTFS의 hyperplane에 포함된다.
- A.1 Lemma 4.1의 증명: causal region C는 ATTFS의 모든 hyperplane의 여집합의 connected component다.
- A.1 Lemma 4.1의 증명: C에서 어떤 affine expression도 부호가 변하지 않으므로 causal set S는 C 전체에서 일정하다.
A.2 Proposition 5.1 증명 · B shallow SNN 보충 세부사항
이 증명은 고정된 causal pattern이 모든 firing time이 입력에 대해 affine인 relatively open convex polyhedron을 정의함을 보인다. 보충 절에서는 이 결과를 shallow network의 bound 및 shared-weight counting 결과와 함께 배치한다.
- A.2 Proposition 5.1 증명: 첫 번째 layer에서는 각 neuron의 causal set을 고정하면 입력 변수에 대한 affine firing times와 linear inequalities가 주어진다.
- A.2 Proposition 5.1 증명: Neuron별 첫 번째 layer 제약을 교집합하면 해당 layer의 causal pattern을 실현하는 입력만 정확히 포함하는 open convex polyhedron이 생성된다.
- A.2 Proposition 5.1 증명: 귀납적으로 affine presynaptic times에 의해 이후의 각 causal condition은 원래 입력에 대한 strict 또는 weak linear inequality가 된다.
- A.2 Proposition 5.1 증명: 따라서 모든 neuron과 layer에 걸쳐 고정된 causal pattern은 모든 firing time이 affine linear인 relatively open convex polyhedron을 정의한다.
- A.2 Proposition 5.1 증명: Causal-pattern region은 equality constraint가 발생하는 boundary set을 제외하면 input space를 partition한다.
- B shallow SNN 보충 세부사항: Shallow-SNN 보충 절에서는 shared weight를 분석하기에 앞서 arrangement 기반 upper bound와 asymptotic lower bound를 증명한다.
- B shallow SNN 보충 세부사항: Shared-weight regime에서 이 보충 절은 shallow spiking network에 대한 exact counting 및 attainability 결과를 확립한다.
B.1 명제 6.2의 증명
이 증명은 R^{d−1}의 N = mκ_d개 hyperplane으로 이루어진 arrangement의 최대 영역 수와 naive causal-set bound를 결합해 서로 다른 causal-set tuple의 수를 bound한다. 그런 다음 더 작은 추정치를 취해 fixed-d, large-m 및 fixed-m, large-d regime에서 결과를 얻는다.
- 명제 6.2의 증명: 서로 다른 causal-set tuple은 arrangement region에서 일정하므로, R^{d−1}의 N = mκ_d개 hyperplane에서 얻는 upper bound와 Proposition 6.1의 naive bound를 결합할 수 있다.관련 defining affine inequality가 tight해질 때에만 causal set이 변한다.
- 명제 6.2의 증명: fixed d, large m에서는 hyperplane-arrangement estimate의 dominant top term을 사용한다.
- 명제 6.2의 증명: fixed m, large d에서는 κ_d = d(2^d−1−1) = Θ(d2^d)와 함께 crude region-count bound를 적용한다.
- 명제 6.2의 증명: arrangement estimate와 naive estimate 중 더 작은 값을 취하면 원하는 bound를 얻는다.
B.2 Proposition 6.3의 증명
각 hidden neuron이 하나의 hyperplane을 유도하는 local domain으로 제한하면 asymptotic lower bound를 확립할 수 있다. 이 domain 안에 bounded region이 포함되는 generic arrangement를 구성하면 서로 다른 causal region이 얻어진다.
- Local hyperplane arrangement: 적절한 local domain의 essentialized space T에서 generic hyperplane arrangement가 만드는 bounded-region count로부터 asymptotic lower bound가 따른다.각 hidden neuron은 이 domain에서 하나의 hyperplane으로 작동한다.
- Local domain: Open box U에서는 처음 d−1개의 input이 t_d보다 먼저 도착하며 서로 2ε 이내로 유지된다.이 box는 |x_i + 1| < ε, 0 < ε < 1 및 x_i = t_i − t_d로 정의된다.
- Causal feasibility: Condition (62)는 t_d 이전에 firing하는 경우 처음 d−1개의 input이 모두 도착할 때까지 기다리도록 강제하므로 causal set을 [d−1]로 고정한다.따라서 U 안에서 hyperplane split 이전의 causal set은 일정하다.
- Causal partition: 각 hidden neuron은 연관된 affine inequality가 positive인지 nonpositive인지에 따라 U를 두 causal region으로 나눈다.두 영역은 inequalities Σ_i w_i^(r)x_i + θ_r > 0 및 Σ_i w_i^(r)x_i + θ_r ≤ 0로 특징지어진다.
- Region counting: 모든 bounded region이 U 안에 놓이도록 hyperplane을 general position으로 선택하면 서로 다른 bounded region이 서로 다른 causal set에 대응한다.offset은 hyperplane이 x_0 = (−1, …, −1) ∈ U에 충분히 가깝게 지나가도록 선택할 수 있다.
B.3 Proposition 6.4의 증명
이 증명은 중첩된 causal-set chain을 세고, 이러한 모든 chain이 shared weights와 ordered thresholds를 사용해 braid cone의 공집합이 아닌 열린 부분집합에서 실현 가능함을 보여 Proposition 6.4를 확립한다.
- Lemma B.1의 증명: Lemma B.1은 neuron들이 positive weights를 공유하고 thresholds가 엄격하게 정렬되어 있을 때 모든 nondecreasing prefix-length sequence를 공집합이 아닌 열린 집합에서 실현한다.고정된 ordering cone 내부에서 해당 causal set은 permutation의 prefix이며, 이를 정의하는 부등식이 strict이므로 perturbation 아래에서도 변하지 않는다.
- Proposition 6.4의 증명: threshold가 증가하면 모든 causal-set sequence는 nested가 되며, Lemma 6.2는 이러한 공집합이 아닌 nested chain을 세어 Proposition 6.4의 upper bound를 얻는다.이 nesting은 Lemma 6.1에서 따르며, counting argument에서는 final set을 고정하고 각 원소에 first entry index를 할당한다.
- Proposition 6.4의 증명: 따라서 upper bound는 tight하다. Lemma 6.2에서 센 모든 nested chain은 shared-weight shallow SNN의 causal-set pattern으로 나타난다.이 construction은 Lemma B.1에 K = d와 N = m을 적용하여 선택한 braid cone의 열린 부분집합에서 주어진 chain을 실현한다.
C Deep SNN 보충 세부사항
Deep-network analysis는 threshold ordering과 shared weights에서 중첩된 prefix causal sets를 도출한 뒤, 이 구조를 사용해 실현 가능한 layer-2 configuration을 열거한다. 또한 two-layer shared-weight architecture가 연속한 구간들을 동일한 비퇴화 output interval에 대응시키는 folding을 수행할 수 없음을 증명한다.
- Deep network 증명: Threshold ordering과 shared incoming weights는 동일 layer의 neuron들 사이에서 중첩된 causal sets와 엄격하게 정렬된 firing times를 강제한다.이 구조적 결과는 deep-network causal-region analysis의 기반이 된다.
- Two-layer 예시: 각 hidden layer에 neuron이 3개인 two-layer network에서는 prefix constraints에 따라 weakly increasing layer-2 label tuple이 정확히 10개 허용된다.가능한 tuple은 {1, 2, 3} 위의 weakly increasing triple이며, (1, 1, 1)에서 (3, 3, 3)까지다.
- Folding 장애: shared-weight (2, m, 2) architecture에서는 정렬된 second-layer thresholds가 relative firing time을 monotone하게 만든다.이 결과는 synaptic delays가 없고 각 layer 내부에서 shared incoming weights가 positive라는 가정에 기반한다.
- Folding 장애: 따라서 second layer는 두 개 이상의 연속한 interval을 하나의 공통 비퇴화 output interval에 affine하고 bijective하게 대응시킬 수 없다.증명은 continuity와 piecewise affinity를 사용해 관련 firing-time map이 nondecreasing임을 보인다.
D 실험 세부사항
실험에서는 주로 random initialization에서 샘플링된 선형 trajectory를 따라 causal-region complexity를 추정하고, architecture, initialization, weights, training이 region count와 accuracy에 미치는 영향을 비교한다. Training 중 SNN은 대응하는 ReLU network보다 더 큰 region complexity를 보이며, initialization은 region count와 CIFAR-10 성능 모두에 큰 영향을 준다.
- 실험 설정: 별도로 명시하지 않는 한 실험은 training 없이 random initialization을 사용하며, 20,001개의 샘플링된 trajectory point에서 region complexity를 추정하고 관측 가능한 region을 20,001개로 제한한다.Forward pass에서는 presynaptic spike time을 정렬하고 최대 d개의 causal prefix를 확인하며, 전체 ANN-style activation vector가 아니라 순서가 정해진 arrival을 사용한다.
- Initialization에서의 region 수: Figure 14는 Init 2가 Init 1보다 더 큰 region count를 생성하며, width와 depth가 증가할수록 count도 증가함을 보여준다. 다만 empirical depth growth는 exponential하지 않다.앞선 layer의 width가 더 클 때 depth의 효과가 더 크게 나타나며, 이는 wide layer가 depth 전반에서 더 높은 차원의 affine image를 보존하는 것과 일치한다.
- Positive weight와 arbitrary weight에서의 region 수: Figure 15는 arbitrary-weight SNN이 일반적으로 Init 1보다 더 많은 region을 실현하며, Init 2가 대부분의 architecture에서 largest count를 달성함을 보여준다.Arbitrary weight에서 count가 더 높아지는 것은 positive-negative cancellation으로 causal-prefix selection이 arrival-time ordering에 더 민감해지기 때문일 수 있으며, 이 mechanism은 아직 체계적으로 연구되지 않았다.
- Training 중 region 수: Figure 16은 초기의 급격한 감소와 부분적 회복에도 불구하고 training 중 SNN의 region complexity가 대응하는 ReLU complexity보다 substantially larger하게 유지됨을 보여준다.Depth-one model은 100 epoch 동안 training하며, width 64, 128, 256에서 10 epoch마다 region complexity를 평가한다.
- Training 중 region 수: Training 중 Init 1과 ReLU는 comparable한 MNIST accuracy를 달성하는 반면, Init 3은 CIFAR-10 SNN performance를 ReLU accuracy에 가깝게 substantially improves한다.따라서 서로 다른 region count에서도 model이 comparable한 MNIST accuracy를 보이지만, initialization은 optimization difficulty와 test performance에 영향을 준다.