본문 바로가기 주메뉴 바로가기
국회도서관 홈으로 정보검색 소장정보 검색

결과 내 검색

동의어 포함

목차보기

Title Page 2

Abstract 7

Contents 9

1. Introduction 15

1.1. Biomarker discovery 15

1.2. Information-theoretic based approach 16

1.3. Graph-based methods in biomarker discovery 17

1.4. Overview 19

2. Feature scoring methods using information-theoretic approaches 21

2.1. Introduction 21

2.2. Feature scoring using reconstruction error as a proxy for mutual information 23

2.2.1. Transforming mutual information into a reconstruction error-based concept 25

2.2.2. Reconstruction error-based feature scoring 27

2.3. Improvements through clustering and limiting bottleneck layer information 30

2.3.1. Simplifying bottleneck layer selection 32

2.3.2. Advanced feature selection via feature-wise clustering 32

2.4. Experiments 33

2.4.1. Performance validation for benchmark datasets 33

2.4.2. Computational cost validation 36

2.4.3. Performance evaluation across varying feature and sample sizes 36

2.4.4. Functional enrichment analysis 36

2.5. Results 37

2.5.1. Performance validation results for benchmark datasets 37

2.5.2. Results of computational cost validation 43

2.5.3. Results of performance evaluation across varying feature and sample sizes 45

2.5.4. Results of functional enrichment analysis: TCGA 47

2.5.5. Results of functional enrichment analysis: ARCHS4 51

2.6. Discussion 53

2.6.1. Discussion: results of TCGA dataset 53

2.6.2. Discussion: results of ARCHS4 dataset 54

2.6.3. Conclusion 55

3. Information theoretic graph-based methods in biomarker discovery 56

3.1. Introduction 56

3.2. Design of a fast algorithm to generate SNP networks 57

3.2.1. Calculation of mutual information through reduction of large-size contingency table 60

3.2.2. Experiments : Simulation Result 63

3.3. Biomarker discovery through conversion from SNP to gene network 66

3.3.1. Construction of SNP epistasis networks using information-theoretic measures 69

3.3.2. Gene-gene interaction network construction from SNP epistasis network 70

3.3.3. Extraction of a statistically significant interaction network 72

3.3.4. Validation through prior knowledge databases 73

3.3.5. Graph refinement and validation using network topology 75

3.3.6. Functional enrichment analysis 83

3.4. Discussion 87

4. Unsupervised feature scoring using reconstruction errors in low-dimensional GNN embeddings 89

4.1. Unsupervised feature scoring that reflects graph characteristics 90

4.2. Performance validation in benchmark datasets 92

4.2.1. Comparative experiments with unsupervised feature selection methods 92

4.2.2. Comparative experiments with the feature importance of a supervised explainable method 96

4.3. Discussion 99

5. Conclusion 100

References 101

List of Tables 14

Table 1. Detailed information of benchmark datasets 35

Table 2. Average accuracy of using 5 to 50 features per method and dataset 39

Table 3. Number of samples per Performance comparison on several benchmark datasets... 42

Table 4. Computational costs comparison of ClearF++ and other feature selection methods 44

Table 5. The top 30 genes with the highest scores obtained from the TCGA dataset 48

Table 6. Significant gene sets of overlap between MSigDB and selected Genes 49

Table 7. Pathway and gene ontology enrichment analysis results using ToppGene on the top... 52

Table 8. Network topologies for gene-gene interaction network measured by mutual... 77

Table 9. Network topologies for gene-gene interaction network measured by information gain 77

Table 10. Network topologies for M.I. and I.G integrated gene-gene interaction network 80

Table 11. Top 10 pathways having the largest gene count from enrichment analysis (p-value 〈... 84

Table 12. Top 10 Gene Ontology terms having the largest gene count from enrichment... 86

List of Figures 12

Figure 1. Overview of the thesis 20

Figure 2. An Overview of supervised feature scoring method using class-wise low-... 24

Figure 3. Simulation results on the relationship between entropy and reconstruction error 26

Figure 4. An illustration of transforming the concept of mutual information into a formula... 26

Figure 5. An illustration of transforming the reconstruction error of the entire dataset into... 29

Figure 6. Simulation results to confirm that our scoring method is suitable for feature... 29

Figure 7. Overview of ClearF++, a supervised feature scoring method that utilizes feature... 31

Figure 8. Cross-validation accuracy in Lung dataset Based on the number of features 38

Figure 9. Cross-validation accuracy in LungDiscrete dataset Based on the number of... 38

Figure 10. Cross-validation accuracy in ProstateGE dataset Based on the number of... 38

Figure 11. Cross-validation accuracy in TCGA dataset Based on the number of features 40

Figure 12. Execution time measurement result in LungDiscrete and ProstatGE dataset for... 44

Figure 13. (A) Performance comparison across varying numbers of features between... 46

Figure 14. Cluster information based on overlap of MsigDB and selected genes 50

Figure 15. Example of making contingency table 59

Figure 16. Example of calculating mutual information 59

Figure 17. Contingency table reduction 61

Figure 18. Compare sample count 61

Figure 19. Construct large-size contingency table 62

Figure 20. Find appropriate large-size contingency table 62

Figure 21. Performance comparison on simulation data 65

Figure 22. Illustration of the overall process of the proposed gene network based... 68

Figure 23. Illustration of the conversion process from a SNP epistasis network to a gene-... 71

Figure 24. Validation process using DisGeNET and GeneMANIA 74

Figure 25. Gene-gene interaction network based on mutual information 78

Figure 26. Top-10 largest components of gene-gene interaction network based on... 79

Figure 27. The largest component of MI and IG integrated gene-gene interaction network 81

Figure 28. Gene-gene interaction network constructed using the GeneMANIA Cytoscape... 82

Figure 29. Overview of unsupervised feature scoring method that reflects graph... 91

Figure 30. Performance comparison of unsupervised feature selection methods on the... 94

Figure 31. Performance comparison of unsupervised feature selection methods on the... 94

Figure 32. Performance comparison of unsupervised feature selection methods on the Cora... 95

Figure 33. Performance comparison of supervised explainable GNN-based methods on the... 97

Figure 34. Performance comparison of supervised explainable GNN-based methods on the... 97

Figure 35. Performance comparison of supervised explainable GNN-based methods on the... 98

초록보기

 Biomarkers are important characteristics that indicate normal biological processes, pathogenic processes, and pharmacological responses, making biomarker development crucial in the fields of medicine and life sciences. Recently, various artificial intelligence models have been developed to identify potential biomarkers, and there is an increasing need for more accurate and reliable methodologies. However, since the biological characteristics of the human body result from complex interactions among multiple features, it is important to employ methodologies that effectively reflect these interactions. Therefore, this thesis proposes biomarker discovery methods that utilize information-theoretic analysis and graph analysis to accurately reflect interactions between features. In the first study, the focus is on developing a stable feature scoring method by replacing the mutual information formula, an information-theoretic relevance measurement method. This approach facilitates faster and more reliable computation of correlations between features and diseases, aiding in biomarker discovery. In the second study, we propose methods for creating and analyzing correlation graphs, using information-theoretic measurements to generate and interpret meaningful graphs. The proposed methods have been validated through comparative experiments in various environments. The first experiment demonstrated that the selected features could be potential biomarker candidates, while the second experiment showed that the generated networks could be useful for biomarker exploration. Subsequent experiments include proposals and analyses of feature selection methods that consider graph structures among samples.