본문 바로가기 주메뉴 바로가기
국회도서관 홈으로 정보검색 소장정보 검색

결과 내 검색

동의어 포함

초록보기

대규모 언어 모델(LLM)의 활용이 확대됨에 따라 사용자 입력을 악용하여 모델의 보안 정책을 우회하는 프롬프트 인젝션 공격이 급증하고 있다.

이를 방지하기 위해 제안된 기존의 방어 기법들은 규칙 기반 탐지로 인해 정확도가 낮거나, 모델 기반 탐지로 높은 연산 비용 문제가 발생하여실시간 서비스 환경에 적용하기 어렵다. 이에 본 연구는 금지어 기반 필터링과 어텐션 패턴 분석(Attention Tracker)를 결합한 고효율 다층 방어탐지 프레임워크를 제안한다. 제안하는 시스템은 1단계에서 공격 키워드를 필터링하여 시스템 부하를 최소화하고, 2단계에서 Focus Score를 통해LLM의 어텐션 가중치 변화를 분석함으로써 문맥 조작 공격을 식별한다. 실제 악성 프롬프트 데이터셋을 이용한 실험 결과, 제안 기법은 기존단일 모델 대비 지연 시간을 50% 이상 단축하였으며, 동시에 약 2배 향상된 탐지 정확도(83%)를 달성하였다. 이를 통해 연산 효율성과 탐지 성능을동시에 확보함으로써, LLM 서비스 환경에서 안정적으로 운용 가능한 실용적인 보안 탐지 체계를 제시하였다.

As the utilization of Large Language Models (LLMs) expands, prompt injection attacks, which manipulate user inputs to bypass modelsecurity policies, have emerged as a significant threat. Existing defense techniques have faced limitations in actual service environmentsdue to the trade-off between the low accuracy of rule-based methods and the high computational costs of model-based detection. Toaddress these challenges, this study proposes a high-efficiency multi-level de4fense framework that combines Banned Terms Filteringwith Attention Pattern Analysis (Attention Tracker). The proposed system minimizes system load by filtering explicit attack keywords inthe first stage and identifies sophisticated attacks that manipulate context by analyzing changes in the LLM’s attention weights usingthe Focus Score in the second stage. Experimental results using a dataset of actual malicious prompts demonstrate that the proposedmethod reduces latency by over 50% and lowers computational costs compared to existing single-model approaches, while achievingapproximately a two-fold increase in detection accuracy (83%). This study is significant in that it presents a practical security detectionsy stem capable of stable operation in LLM service environments by effectively securing both computational efficiency and detectionperformance.

권호기사

권호기사 목록 테이블로 기사명, 저자명, 페이지, 원문, 기사목차 순으로 되어있습니다.
기사명 저자명 페이지 원문 목차
CBDC 시스템 아키텍처의 R-ABAC 접근 권한 관리 모델 연구 = A study on the R-ABAC access privilege management model for CBDC system architecture 김지민, 박건우, 김솔리 p. 273-280
어텐션 패턴 분석을 활용한 다층 프롬프트 인젝션 탐지 프레임워크 = Multi-level prompt injection detection framework using attention pattern analysis 정선우, 김남령, 이일구 p. 281-289
정량적 동작 패턴 기반 RTOS 하드웨어 무결성 위협 탐지 = RTOS hardware integrity threat detection based on quantitative behavior patterns 하영빈, 안성규, 박기웅 p. 290-297
안드로이드에서 안티-포렌식 대응을 위한 커널 수준의 실시간 타임스탬프 변조 탐지 기법 = Kernel-level real-time detection of timestamp manipulation on Android for anti-forensics resistance 안균승, 안석현, 조성제 p. 298-305
Analyzing object-scale dependency of dual-branch sigmoid CAM for weakly supervised object localization = 약지도 객체 위치추정을 위한 이중 분기 시그모이드 CAM의 객체 크기 의존성 분석 Eunhyun Ryu, Hyewon Joo, Junhyug Noh p. 306-314
건강검진정보를 이용한 다층신경망 기반 성별 생물학적 나이 예측 연구 = A multilayer neural network–based study on sex-specific biological age prediction using health checkup data 이상호, 송주영, 최경식, 전지영, 송태민 p. 315-323
Informativeness-aware layer freezing and sample replay for efficient online continual object detection = 정보성 인지 계층 동결 및 샘플 리플레이 기반 온라인 지속 객체 탐지 Taeheon Kim, Minjae Lee, Jonghyun Park, Chaeeun Lee, Changdae Lee, Mingi Kim, Dongseok Lee, Jonghyun Choi p. 324-333
그래디언트 감쇠를 통한 양자화 인지 학습의 진동 억제 = Suppressing oscillations in quantization-aware training via gradient attenuation 이지호, 이형섭, 강우철 p. 334-340
질문 변형 및 표 이미지 병합 기법을 통한 표 데이터 이해 성능 향상 = Enhancing table understanding performance through question variation and table image merging techniques 신미르, 신유현 p. 341-348
원샷 비율 탐색 및 인-트레이닝 희소도 스케줄링 기반의 적응형 필터 프루닝 기법 = Adaptive filter pruning via one-shot ratio search and in-training sparsity scheduling 이현수, 문용혁 p. 349-358