본문 바로가기 주메뉴 바로가기
국회도서관 홈으로 정보검색 소장정보 검색

결과 내 검색

동의어 포함

목차보기

Title Page 2

Contents 5

Abstract 11

Chapter 1. Introduction 13

Chapter 2. Related Works 17

2.1. Subgraph Federated Learning 17

2.2. Curriculum Graph Learning 18

Chapter 3. Preliminaries 20

3.1. Graph Neural Networks 20

3.2. Personalized Subgraph FL Optimization 21

Chapter 4. Methods 22

4.1. Local Training Stage 22

4.1.1. Incremental Edge Selection (IES) 22

4.1.2. Local Model Optimization 24

4.2. Server Aggregation Stage 25

4.2.1. Client Similarity Estimation via Node Embedding Distributions 26

4.2.2. Personalized Parameter Aggregation 28

4.3. Complexity Analysis 29

4.4. Theoretical Analysis 29

Chapter 5. Experiments 31

5.1. Experimental Setup 31

5.1.1. Datasets 31

5.1.2. Baselines 31

5.1.3. Hyperparameters 32

5.2. Main Results 32

5.2.1. Performance and Effectiveness 32

5.2.2. Impact of CL on Server Aggregation 35

5.3. Ablation Study 36

5.3.1. Hyperparameter Analysis on Curriculum 36

5.3.2. Regularization in Local Training Stage 37

5.3.3. Effectiveness of Automatic CL Strategy 37

5.3.4. Edge Filtering Approaches for Client Similarity Estimation 38

5.3.5. Varying Scaling Factor τ 39

5.3.6. Results on Louvain Partitioning 40

5.3.7. Varying the Random Graph Model 41

5.3.8. Assessing the Fidelity of Client Similarity Estimation 42

Chapter 6. Discussion 44

Chapter 7. Conclusion 45

References 46

Chapter A. Appendix 56

A.1. Server Aggregation Stage Algorithm 56

A.2. Detailed Theoretical Analysis 57

A.2.1. Preliminaries and Assumptions 57

A.2.2. Community Preservation 58

A.2.3. Generalization Bound and Overfitting Control 60

A.3. Dataset Descriptions 62

A.4. Baselines 63

A.4.1. FedAvgCL 63

A.4.2. FedGNN 63

A.5. Auxiliary Experiments 64

A.6. Computing Resources 64

Abstract 65

List of Tables 10

Table 5.1. Node classification performance on FL frameworks over three different numbers... 33

Table 5.2. Performance gain from CL and proximal term 36

Table 5.3. Performance comparison according to CL strategies. Each strategy except ours... 37

Table 5.4. Performance variation according to the scaling factor. "adaptive" refers to the... 39

Table 5.5. Performance on subgraphs generated via the Louvain algorithm 40

Table 5.6. Performance of CUFL when the reference graph is generated by different random... 41

Table A.1. Dataset statistics. The number of nodes, edges, and classes for each setting is... 62

List of Figures 9

Figure 1.1. Training trends of three FL frameworks on Cora with 10 clients. (A) Cross-... 14

Figure 4.1. Overview of the CUFL framework for Client 1. (A) Incremental Edge... 23

Figure 4.2. Edge-wise bin-match ratios with respect to Client 1 on the Cora dataset. The... 27

Figure 5.1. Accuracy curves for six FL frameworks. FedSpray training is stopped after 100... 34

Figure 5.2. Evolution of weight proportions of data-similar clients during server aggrega-... 34

Figure 5.3. Performance of local training under different hyperparameter settings for learn-... 36

Figure 5.4. Heatmaps of estimated client similarity on Cora with 10 clients under the node... 38

Figure 5.5. Evolution of weight proportions of data-similar clients during server aggre-... 40

Figure 5.6. Actual versus estimated subgraph similarity on Cora with 10 clients in the... 42

Figure A.1. Class distribution of datasets. Darker color indicates that more nodes belong... 62

초록보기

 서브그래프 연합 학습 (Federated Learning, FL)은 여러 개의 비공개 서브그래프에 분산된 그래프 신경망 (Graph Neural Networks, GNN) 을 공동 학습하지만, 서브그래프 간 데이터 이질성 (Data Heterogeneity) 으로 인해 학습이 어렵다. 이 이질성을 완화하기 위해, 가중치 기반 모델 집계는 각 클라이언트의 현재 모델 상태에서 추정한 서브그래프 특성이 비슷한 클라이언트일수록 더 큰 가중치를 부여해 로컬 GNN을 개인화한다. 구체적으로, 서버는 원본 서브그래프 데이터를 공유하지 않고도 개인정보를 보호하는 모델 지표를 비교해 클라이언트 유사도 행렬을 만들고, 유사도가 높은 클라이언트끼리 서로의 업데이트에 더 큰 비중을 두도록 집계 규모를 조절한다. 그러나 클라이언트가 보유한 서브그래프는 희소하고 편향돼 있어 빠른 과적합을 유발하고, 그 결과 유사도 행렬이 정체되거나 붕괴될 수 있다. 이 경우 각 클라이언트는 다양한 지식을 흡수하지 못하고 자기 편향만 강화하게 되어 집계 효과가 사라진다.

이를 해결하기 위해 본 논문은 Curriculum guided personalized sUbgraph Federated Learning (CUFL)을 제안한다. 클라이언트 측에서는 커리큘럼 학습 (Curriculum Learning, CL) 을 도입해 재구성 점수에 따라 학습에 사용할 엣지를 자동으로 선택한다. 먼저 쉬운 범용 구조를 GNN 에 노출하고 이후 점진적으로 어려운, 클라이언트 특화 구조를 제공함으로써 초기 과적합을 억제하고 점진적 개인화를 가능하게 한다. 이렇게 개인화 정도를 조절함으로써, 서버 집계 역시 범용 지식을 주고받는 단계에서 클라이언트별 지식을 전파하는 단계로 자연스럽게 전환된다. 또한 CUFL은 무작위 참조 그래프를 재구성해 얻은 미세한 구조적 지표로 클라이언트 유사도를 추정함으로써 가중치 집계의 정밀도를 높였다. 6개 벤치마크 데이터셋에서의 광범위한 실험은 CUFL이 기존 방법보다 우수한 성능을 달성함을 확인한다. 코드는 다음 링크에서 제공된다 https://github.com/Kang-Min-Ku/Curriculum-Guided-FL.git.