SOP(표준작업지침서), 배치 제조기록서, 기술 보고서, 규제 제출 문서 등 사내에 축적된 수많은 문서가 있고, 이를 기반으로 구동되는 AI 기반 검색 시스템을 구축하고자 합니다. 하지만 회사의 기밀 문서나 고유 데이터를 임베딩 생성을 위해 서드파티 외부 API로 전송하는 방식은 원치 않을 것입니다. 모든 바이트가 사내 인프라 내에 그대로 머물고, 임베딩이 로컬 환경에서 생성되며, 전체 검색 파이프라인을 완벽하게 통제할 수 있는 시스템이 필요합니다.
이를 실현하는 최적의 아키텍처가 바로 하이브리드 RAG(Hybrid RAG)입니다. 고밀도 벡터 검색(시맨틱 의미)과 희소 BM25 키워드 매칭(정확한 용어)을 결합하고, 그 결과를 융합(fuse)한 후, 질문과 연관된 청크(chunk)만을 선별하여 프론티어 모델(Frontier Model)에 전달함으로써 최종 답변을 생성합니다. 본 튜토리얼에서는 Docker Compose를 활용하여 단 하나의 명령어로 실행되는 4개의 컨테이너 스택 전체를 구축합니다. 이 레퍼런스 구현의 전체 소스 코드는 GitHub hybrid-rag-docker-example 리포지토리에서 확인할 수 있습니다.
하이브리드 검색이 중요한 이유 (Why hybrid matters)
순수 벡터 검색(Pure vector search)은 품질 및 규제 문서에서 종종 실패합니다. “SOP-QA-007”과 “SOP-QA-070”은 거의 동일한 임베딩 벡터를 생성합니다. “USP <85>”(세균성 내독소)와 “USP <87>”(생물학적 반응성)은 문맥상 매우 유사하게 기술되지만 규제 관점에서는 완전히 다른 규격입니다. 벡터 검색이 이를 혼동하는 이유는 임베딩이 정확한 문자열이 아니라 의미(semantic meaning)를 포착하기 때문입니다.
반대로 순수 BM25 키워드 검색은 정반대 방향에서 실패합니다. “보관 온도가 임계값 아래로 떨어지면 어떻게 됩니까(what happens if storage temperature drops below threshold)“라고 질문하면, “열 센서 측정값이 최소값 미만으로 떨어질 때 콜드룸 일탈 이벤트 발생(cold room deviation event triggered when thermal reading falls beneath the minimum)“이라고 적힌 청크를 놓치게 됩니다. 동일한 의미이지만 사용된 단어가 완전히 다르기 때문입니다.
하이브리드 검색은 이 두 가지를 모두 잡아냅니다. 고밀도 검색(Dense retrieval)은 개념적 일치를 처리하고, 희소 검색(Sparse retrieval)은 정확한 용어 일치를 처리합니다. 그리고 상호 순위 융합(Reciprocal Rank Fusion, RRF)은 두 개의 순위화된 목록을 단일 결과 세트로 병합하여, 어느 한 방식만을 단독으로 사용할 때보다 훨씬 뛰어난 검색 품질을 제공합니다.
전체 아키텍처 (The architecture)
┌────────────────────────── Docker Compose Network ───────────────────────────┐
│ │
│ ┌──────────────┐ ┌──────────────────┐ ┌──────────────────┐ │
│ │ │ │ │ │ │ │
│ │ RustFS │◀──────│ RAG API │──────▶│ Ollama │ │
│ │ (S3 Store) │ │ (Orchestrator) │ │ qwen3-embedding │ │
│ │ :9000 │ │ :8000 │ │ :11434 │ │
│ │ │ │ │ │ │ │
│ └──────────────┘ │ ┌────────────┐ │ └──────┬───────────┘ │
│ │ │ LanceDB │ │ │ │
│ │ │ (Vector + │ │ │ │
│ │ │ BM25 FTS) │ │ │ │
│ │ └───────┬────┘ │ ┌──────▼───────────┐ │
│ └──────────┴───────┘ │ Frontier Model │ │
│ │ │ (OpenAI-compat) │ │
│ │ │ via Ollama │ │
│ ┌─────▼────┐ └──────────────────┘ │
│ │ Client │ │
│ └──────────┘ │
└─────────────────────────────────────────────────────────────────────────────┘
동작 흐름 (The flow):
- Rust로 작성된 S3 호환 객체 스토리지인 RustFS에 문서를 업로드합니다.
- RAG API가 파일을 다운로드하고 파싱한 뒤, 일부가 겹치는(overlapping) 청크 단위로 분할합니다.
- Ollama가
qwen3-embedding:8b를 사용하여 4096차원 임베딩을 생성합니다. - LanceDB가 고밀도 벡터와 Tantivy 기반 전문 검색(Full-Text Search) 인덱스 모두와 함께 청크를 저장합니다.
- 검색 질의 시: LanceDB가 벡터 검색과 BM25 키워드 검색을 동시에 수행합니다.
- 상호 순위 융합(RRF) 알고리즘으로 검색 결과를 병합합니다.
- 조합된 컨텍스트가 OpenAI 호환 엔드포인트를 통해 프론티어 모델로 전달됩니다.
- 프론티어 모델이 원본 문서에 근거한(grounded) 신뢰도 높은 답변을 생성합니다.
프로젝트 디렉토리 구조 (Project structure)
hybrid-rag/
├── docker-compose.yml
├── .env
├── orchestrator/
│ ├── Dockerfile
│ ├── requirements.txt
│ ├── app/
│ │ ├── __init__.py
│ │ ├── main.py
│ │ ├── config.py
│ │ ├── embedding.py
│ │ ├── document_processor.py
│ │ ├── vector_store.py
│ │ ├── hybrid_search.py
│ │ └── llm_client.py
│ └── scripts/
│ └── init-ollama.sh
└── data/
└── sample/
└── sample.txt
단계 1: Docker Compose 구성 (Step 1: Docker Compose)
모든 구성 요소를 하나로 묶어주는 단일 설정 파일입니다. 4개의 서비스, 1개의 네트워크, 그리고 모든 상태를 보존하기 위한 영구 볼륨(persistent volume)으로 구성됩니다.
version: '3.8'
services:
rustfs:
image: rustfs/rustfs:latest
container_name: rustfs
ports:
- "9000:9000"
- "9001:9001"
environment:
- RUSTFS_ROOT_USER=admin
- RUSTFS_ROOT_PASSWORD=admin123456
volumes:
- rustfs_data:/data
command: server /data --console-address ":9001"
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:9000/minio/health/live"]
interval: 15s
timeout: 10s
retries: 5
start_period: 30s
restart: unless-stopped
ollama:
image: ollama/ollama:latest
container_name: ollama
ports:
- "11434:11434"
volumes:
- ollama_data:/root/.ollama
- ./scripts/init-ollama.sh:/docker-entrypoint-init.d/init-ollama.sh:ro
# Uncomment for NVIDIA GPU:
# deploy:
# resources:
# reservations:
# devices:
# - driver: nvidia
# count: all
# capabilities: [gpu]
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:11434/api/tags"]
interval: 30s
timeout: 10s
retries: 10
start_period: 60s
restart: unless-stopped
rag-orchestrator:
build:
context: ./orchestrator
dockerfile: Dockerfile
container_name: rag-orchestrator
ports:
- "8000:8000"
environment:
- RUSTFS_ENDPOINT=rustfs:9000
- RUSTFS_ACCESS_KEY=admin
- RUSTFS_SECRET_KEY=admin123456
- RUSTFS_BUCKET=documents
- OLLAMA_HOST=http://ollama:11434
- EMBEDDING_MODEL=qwen3-embedding:8b
- FRONTIER_MODEL_URL=http://ollama:11434/v1
- FRONTIER_MODEL_NAME=qwen3:32b
- LANCEDB_URI=/app/data/lancedb
- CHUNK_SIZE=512
- CHUNK_OVERLAP=50
- TOP_K=5
- DENSE_WEIGHT=0.7
volumes:
- lancedb_data:/app/data/lancedb
depends_on:
rustfs:
condition: service_healthy
ollama:
condition: service_healthy
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 15s
timeout: 5s
retries: 5
start_period: 30s
restart: unless-stopped
volumes:
rustfs_data:
ollama_data:
lancedb_data:
RustFS는 UID 10001 권한으로 실행됩니다. 데이터 볼륨에서 권한 오류가 발생하는 경우 다음 명령을 실행하십시오:
sudo chown -R 10001:10001 ./data/rustfs
단계 2: Ollama 초기화 스크립트 (Step 2: Ollama initialization script)
오케스트레이터에는 임베딩을 위한 qwen3-embedding:8b와 텍스트 생성을 위한 qwen3:32b 두 가지 모델이 필요합니다. 이 스크립트는 Ollama 컨테이너가 처음 구동될 때 내부에서 실행되어 두 모델을 자동으로 다운로드(pull)합니다.
#!/bin/bash
set -e
# Wait for Ollama API
MAX_RETRIES=60
RETRY_COUNT=0
while ! curl -sf http://localhost:11434/api/tags > /dev/null 2>&1; do
RETRY_COUNT=$((RETRY_COUNT + 1))
if [ $RETRY_COUNT -ge $MAX_RETRIES ]; then
echo "ERROR: Ollama API not ready"
exit 1
fi
sleep 5
done
# Pull embedding model
if ! ollama list | grep -q "^qwen3-embedding:8b"; then
echo "Pulling qwen3-embedding:8b..."
ollama pull qwen3-embedding:8b
fi
# Pull frontier model
if ! ollama list | grep -q "^qwen3:32b"; then
echo "Pulling qwen3:32b..."
ollama pull qwen3:32b
fi
echo "Models ready."
ollama list
실행 권한을 부여합니다: chmod +x scripts/init-ollama.sh
리소스 사용량을 낮추려면 qwen3:32b 대신 qwen3:8b를 사용할 수 있으며, 전체 아키텍처는 완전히 동일하게 유지됩니다.
단계 3: 오케스트레이터 Dockerfile (Step 3: The orchestrator Dockerfile)
FROM python:3.11-slim-bookworm
WORKDIR /app
RUN apt-get update && apt-get install -y --no-install-recommends \
build-essential curl && \
rm -rf /var/lib/apt/lists/*
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
EXPOSE 8000
CMD ["uvicorn", "app.main:app", "--host", "0.0.0.0", "--port", "8000", "--workers", "1"]
단계 4: Python 의존성 패키지 (Step 4: Python dependencies)
fastapi==0.115.6
uvicorn[standard]==0.34.0
pydantic==2.10.4
boto3==1.35.99
lancedb==0.20.0
pandas==2.2.3
pyarrow==18.1.0
requests==2.32.3
pypdf==5.1.0
python-multipart==0.0.20
numpy==2.2.1
python-dotenv==1.0.1
tqdm==4.67.1
단계 5: 애플리케이션 환경 설정 (Step 5: Configuration)
# app/config.py
import os
from dataclasses import dataclass, field
@dataclass
class Config:
# RustFS
rustfs_endpoint: str = os.getenv("RUSTFS_ENDPOINT", "rustfs:9000")
rustfs_access_key: str = os.getenv("RUSTFS_ACCESS_KEY", "admin")
rustfs_secret_key: str = os.getenv("RUSTFS_SECRET_KEY", "admin123456")
rustfs_bucket: str = os.getenv("RUSTFS_BUCKET", "documents")
# Ollama
ollama_host: str = os.getenv("OLLAMA_HOST", "http://ollama:11434")
embedding_model: str = os.getenv("EMBEDDING_MODEL", "qwen3-embedding:8b")
# Frontier Model (OpenAI-compatible)
frontier_url: str = os.getenv("FRONTIER_MODEL_URL", "http://ollama:11434/v1")
frontier_model: str = os.getenv("FRONTIER_MODEL_NAME", "qwen3:32b")
frontier_api_key: str = os.getenv("FRONTIER_API_KEY", "ollama")
# LanceDB
lancedb_uri: str = os.getenv("LANCEDB_URI", "/app/data/lancedb")
table_name: str = "documents"
# Chunking
chunk_size: int = int(os.getenv("CHUNK_SIZE", "512"))
chunk_overlap: int = int(os.getenv("CHUNK_OVERLAP", "50"))
# Retrieval
top_k: int = int(os.getenv("TOP_K", "5"))
dense_weight: float = float(os.getenv("DENSE_WEIGHT", "0.7"))
rrf_k: int = 60
settings = Config()
단계 6: 문서 처리 및 청킹 (Step 6: Document processing)
PDF, DOCX, TXT, Markdown 문서를 처리합니다. 문단 및 문장 경계를 존중하는 재귀적 문자 분할(recursive character splitting) 방식을 사용하여 텍스트를 청킹합니다.
# app/document_processor.py
import io
import re
import pypdf
def extract_text(file_content: bytes, filename: str) -> str:
ext = filename.rsplit(".", 1)[-1].lower()
if ext == "pdf":
reader = pypdf.PdfReader(io.BytesIO(file_content))
return "\n".join(page.extract_text() or "" for page in reader.pages)
elif ext == "docx":
from docx import Document
doc = Document(io.BytesIO(file_content))
return "\n".join(p.text for p in doc.paragraphs)
else:
return file_content.decode("utf-8", errors="replace")
def chunk_text(text: str, chunk_size: int = 512, overlap: int = 50) -> list[str]:
text = text.strip()
if len(text) <= chunk_size:
return [text]
separators = ["\n\n", "\n", ". ", " ", ""]
return _split(text, separators, chunk_size, overlap)
def _split(text, separators, chunk_size, overlap):
if len(text) <= chunk_size:
return [text] if text.strip() else []
sep = next((s for s in separators if s in text), separators[-1])
parts = text.split(sep)
chunks, current = [], ""
for part in parts:
candidate = (current + sep + part) if current else part
if len(candidate) <= chunk_size:
current = candidate
else:
if current:
chunks.append(current)
if len(part) > chunk_size and len(separators) > 1:
chunks.extend(_split(part, separators[1:], chunk_size, overlap))
current = ""
else:
current = part
if current and current.strip():
chunks.append(current)
if overlap > 0 and len(chunks) > 1:
overlapped = [chunks[0]]
for i in range(1, len(chunks)):
overlapped.append(chunks[i-1][-overlap:] + " " + chunks[i])
chunks = overlapped
return [c.strip() for c in chunks if c.strip()]
단계 7: 임베딩 생성 (Step 7: Embedding generation)
Ollama의 임베딩 API를 호출합니다. qwen3-embedding:8b 모델은 4096차원 벡터를 생성하며, 이는 의미론적으로 유사하지만 기능적으로 상이한 문서 섹션들을 정밀하게 구별해낼 수 있을 만큼 높은 해상도를 제공합니다.
# app/embedding.py
import requests
from app.config import settings
def get_embedding(text: str) -> list[float]:
resp = requests.post(
f"{settings.ollama_host}/api/embeddings",
json={"model": settings.embedding_model, "prompt": text.strip()},
timeout=120
)
resp.raise_for_status()
return resp.json()["embedding"]
def get_embeddings_batch(texts: list[str]) -> list[list[float]]:
embeddings = []
with requests.Session() as session:
for text in texts:
resp = session.post(
f"{settings.ollama_host}/api/embeddings",
json={"model": settings.embedding_model, "prompt": text.strip()},
timeout=120
)
resp.raise_for_status()
embeddings.append(resp.json()["embedding"])
return embeddings
단계 8: LanceDB 벡터 스토어 (Step 8: LanceDB vector store)
LanceDB는 단일 테이블 내에 고밀도 벡터와 Tantivy 기반 전문 검색 인덱스를 모두 저장합니다. 별도의 외부 의존성 없이 오케스트레이터 프로세스 내에 임베디드(embedded) 형태로 실행됩니다.
# app/vector_store.py
import logging
import pandas as pd
import pyarrow as pa
import lancedb
from app.config import settings
logger = logging.getLogger(__name__)
_db = None
_table = None
def get_table():
global _db, _table
if _table is None:
_db = lancedb.connect(settings.lancedb_uri)
if settings.table_name in _db.table_names():
_table = _db.open_table(settings.table_name)
else:
_table = None
return _table
def init_table(dimension: int):
global _db, _table
_db = lancedb.connect(settings.lancedb_uri)
if settings.table_name in _db.table_names():
_table = _db.open_table(settings.table_name)
logger.info("Opened table '%s' (%d rows)", settings.table_name, _table.count_rows())
else:
schema = pa.schema([
pa.field("chunk_id", pa.string()),
pa.field("text", pa.string()),
pa.field("source", pa.string()),
pa.field("vector", pa.list_(pa.float32(), list_size=dimension)),
])
_table = _db.create_table(settings.table_name, schema=schema)
logger.info("Created table '%s' with dim=%d", settings.table_name, dimension)
return _table
def add_records(records: list[dict]):
global _table
if _table is None:
raise RuntimeError("Table not initialized")
_table.add(records)
단계 9: 상호 순위 융합(RRF) 기반 하이브리드 검색 (Step 9: Hybrid search with Reciprocal Rank Fusion)
하이브리드 검색의 핵심 로직입니다. 고밀도 벡터 유사도 검색과 희소 BM25 키워드 매칭이라는 두 가지 검색이 병렬로 실행된 뒤, RRF(Reciprocal Rank Fusion)를 통해 결과가 하나로 융합됩니다.
# app/hybrid_search.py
import numpy as np
from rank_bm25 import BM25Okapi
from app.config import settings
from app.vector_store import get_table
from app.embedding import get_embeddings_batch
from app.vector_store import get_table
class HybridSearchEngine:
def __init__(self):
self.all_chunks = []
self.bm25 = None
self.ready = False
def rebuild(self):
table = get_table()
if table is None:
return
df = table.to_pandas()
self.all_chunks = df.to_dict("records")
if self.all_chunks:
tokenized = [r["text"].lower().split() for r in self.all_chunks]
self.bm25 = BM25Okapi(tokenized)
self.ready = True
def vector_search(self, query_embedding, limit):
table = get_table()
return table.search(query_embedding).metric("cosine").limit(limit).to_list()
def keyword_search(self, query, limit):
if self.bm25 is None:
return []
scores = self.bm25.get_scores(query.lower().split())
top_idx = np.argsort(scores)[::-1][:limit]
return [{**self.all_chunks[i], "_bm25_score": float(scores[i])}
for i in top_idx if scores[i] > 0]
def hybrid_search(self, query, top_k=None):
top_k = top_k or settings.top_k
fetch_k = top_k * 3
query_embedding = get_embeddings_batch([query])[0]
vec_results = self.vector_search(query_embedding, fetch_k)
kw_results = self.keyword_search(query, fetch_k)
# RRF fusion
rrf_scores = {}
id_to_result = {}
for rank, r in enumerate(vec_results):
cid = r["chunk_id"]
rrf_scores[cid] = rrf_scores.get(cid, 0.0) + 1.0 / (settings.rrf_k + rank + 1)
id_to_result[cid] = r
for rank, r in enumerate(kw_results):
cid = r["chunk_id"]
rrf_scores[cid] = rrf_scores.get(cid, 0.0) + 1.0 / (settings.rrf_k + rank + 1)
if cid not in id_to_result:
id_to_result[cid] = r
sorted_ids = sorted(rrf_scores, key=rrf_scores.get, reverse=True)[:top_k]
final = []
for cid in sorted_ids:
entry = id_to_result[cid].copy()
entry["_rrf_score"] = round(rrf_scores[cid], 6)
final.append(entry)
return final
engine = HybridSearchEngine()
RRF 동작 원리 (How RRF works):
RRF_score(d) = Σ 1 / (k + rank_i(d))
여기서 k = 60(평활화 상수, smoothing constant)이며, rank_i(d)는 각 결과 목록에서 해당 문서의 순위 위치(1부터 시작)입니다. 예를 들어 벡터 검색에서 3위, 키워드 검색에서 5위를 차지한 청크의 점수는 다음과 같이 계산됩니다:
score = 1/(60+3) + 1/(60+5) = 0.01587 + 0.01538 = 0.03125
이 방식은 두 검색 결과 목록에 모두 등장하는 문서에 자연스럽게 가중치를 부여하며, 이는 하이브리드 검색이 추구하는 목적에 정확히 부합합니다.
단계 10: LLM 클라이언트 (Step 10: LLM client)
프론티어 모델은 OpenAI 호환 엔드포인트를 통해 호출됩니다. 따라서 애플리케이션 코드를 전혀 수정하지 않고도 Ollama, vLLM, LiteLLM, OpenAI, OpenRouter 또는 기타 호환 프록시로 손쉽게 전환할 수 있습니다.
# app/llm_client.py
from openai import OpenAI
from app.config import settings
client = OpenAI(
base_url=settings.frontier_url,
api_key=settings.frontier_api_key,
)
def generate_answer(query: str, context_chunks: list[str]) -> str:
context = "\n\n---\n\n".join(
f"[Source {i+1}]: {chunk}" for i, chunk in enumerate(context_chunks)
)
system_prompt = (
"You are a precise assistant. Answer using ONLY the provided context. "
"If the context doesn't contain enough information, say so clearly. "
"Cite sources by their [number] reference."
)
user_prompt = f"Context:\n\n{context}\n\n---\n\nQuestion: {query}\n\nAnswer:"
response = client.chat.completions.create(
model=settings.frontier_model,
messages=[
{"role": "system", "content": system_prompt},
{"role": "user", "content": user_prompt},
],
temperature=0.2,
max_tokens=2048,
)
return response.choices[0].message.content
단계 11: FastAPI 메인 애플리케이션 (Step 11: FastAPI application)
메인 애플리케이션은 파일 업로드, 수집(ingestion), 하이브리드 검색 질의, 헬스 체크 등 모든 기능을 유기적으로 연결합니다.
# app/main.py
import time
import logging
from contextlib import asynccontextmanager
from fastapi import FastAPI, UploadFile, File, HTTPException
from pydantic import BaseModel
import boto3
from botocore.client import Config as BotoConfig
from app.config import settings
from app.document_processor import extract_text, chunk_text
from app.embedding import get_embeddings_batch, get_embedding
from app.vector_store import init_table, add_records
from app.hybrid_search import engine
from app.llm_client import generate_answer
logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)
s3 = boto3.client(
"s3",
endpoint_url=f"http://{settings.rustfs_endpoint}",
aws_access_key_id=settings.rustfs_access_key,
aws_secret_access_key=settings.rustfs_secret_key,
config=BotoConfig(signature_version="s3v4"),
region_name="us-east-1",
)
@asynccontextmanager
async def lifespan(app: FastAPI):
# Ensure bucket exists
try:
s3.head_bucket(Bucket=settings.rustfs_bucket)
except Exception:
s3.create_bucket(Bucket=settings.rustfs_bucket)
# Probe embedding dimension
dim = len(get_embedding("dimension probe"))
init_table(dim)
engine.rebuild()
logger.info("Ready. %d chunks indexed. Embedding dim=%d.", engine.chunk_count, dim)
yield
app = FastAPI(title="Hybrid RAG Orchestrator", lifespan=lifespan)
class QueryRequest(BaseModel):
query: str
top_k: int = settings.top_k
class QueryResponse(BaseModel):
answer: str
sources: list[dict]
timings: dict
@app.get("/health")
async def health():
return {
"status": "ok",
"documents": engine.document_count,
"chunks": engine.chunk_count,
"embedding_model": settings.embedding_model,
"llm_model": settings.frontier_model,
}
@app.post("/api/upload")
async def upload_file(file: UploadFile = File(...)):
data = await file.read()
s3.put_object(Bucket=settings.rustfs_bucket, Key=file.filename, Body=data)
return {"filename": file.filename, "size": len(data)}
@app.post("/api/ingest")
async def ingest():
objects = s3.list_objects_v2(Bucket=settings.rustfs_bucket).get("Contents", [])
total_chunks = 0
files_processed = 0
for obj in objects:
key = obj["Key"]
ext = key.rsplit(".", 1)[-1].lower()
if ext not in {"pdf", "txt", "md", "docx"}:
continue
raw = s3.get_object(Bucket=settings.rustfs_bucket, Key=key)["Body"].read()
text = extract_text(raw, key)
chunks = chunk_text(text, settings.chunk_size, settings.chunk_overlap)
if not chunks:
continue
embeddings = get_embeddings_batch(chunks)
records = [
{"chunk_id": f"{key}#{i}", "text": c, "source": key, "vector": e}
for i, (c, e) in enumerate(zip(chunks, embeddings))
]
add_records(records)
total_chunks += len(records)
files_processed += 1
engine.rebuild()
return {"files": files_processed, "chunks": total_chunks}
@app.post("/api/query", response_model=QueryResponse)
async def query(req: QueryRequest):
if engine.chunk_count == 0:
raise HTTPException(400, "No documents ingested. Upload and run /api/ingest first.")
t0 = time.time()
results = engine.hybrid_search(req.query, req.top_k)
search_ms = round((time.time() - t0) * 1000)
context_chunks = [r["text"] for r in results]
sources = [
{"source": r["source"], "text": r["text"][:200], "rrf_score": r.get("_rrf_score", 0)}
for r in results
]
t0 = time.time()
answer = generate_answer(req.query, context_chunks)
llm_ms = round((time.time() - t0) * 1000)
return QueryResponse(
answer=answer,
sources=sources,
timings={"search_ms": search_ms, "llm_ms": llm_ms, "total_ms": search_ms + llm_ms},
)
단계 12: 실행 및 확인 (Step 12: Run it)
# 전체 서비스 실행
docker compose up -d --build
# 오케스트레이터 로그 확인 ("Ready" 메시지 대기)
docker compose logs -f rag-orchestrator
# 문서 업로드
curl -X POST http://localhost:8000/api/upload -F "file=@sop-cleanroom-01.txt"
# 수집 트리거 (청킹, 임베딩, 인덱싱)
curl -X POST http://localhost:8000/api/ingest
# 질의 테스트
curl -X POST http://localhost:8000/api/query \
-H "Content-Type: application/json" \
-d '{"query": "What happens if CR-MODULE-99 reports a low temperature?"}'
단계 13: 테스트 문서를 통한 검증 (Step 13: Verify with a test document)
sop-cleanroom-01.txt 파일을 생성합니다:
SOP-CLEANROOM-01: Environmental Monitoring Parameters.
All cold room storage arrays designated under tier-1 biological compliance
must operate continuously at a threshold setting of 4.0°C. If sensor reading
array CR-MODULE-99 drops below 2.0°C, an absolute deviation event is
registered, and quality control management must be notified immediately.
문서를 업로드하고 수집(ingest)한 후, 정확히 일치하는 용어 질의와 시맨틱 표현 질의를 각각 실행해 봅니다:
# 정확한 키워드 매칭 테스트
curl -X POST http://localhost:8000/api/query \
-H "Content-Type: application/json" \
-d '{"query": "What is the threshold for CR-MODULE-99?"}'
# 시맨틱 유사도 매칭 테스트
curl -X POST http://localhost:8000/api/query \
-H "Content-Type: application/json" \
-d '{"query": "What happens if storage temperature drops too low?"}'
첫 번째 질의는 정확한 식별자인 “CR-MODULE-99”를 찾기 위해 BM25 키워드 매칭에 크게 의존합니다. 두 번째 질의는 “온도가 너무 낮게 떨어지면(drops too low)“이라는 표현이 “2.0°C 미만으로 떨어짐(drops below 2.0°C)“과 동일한 의미임을 이해하기 위해 벡터 유사도에 의존합니다. 하이브리드 검색은 이 두 가지 상황을 모두 완벽하게 처리합니다.
검색 가중치 튜닝 (Tuning the retrieval weights)
DENSE_WEIGHT 환경 변수를 통해 가중치 균형을 조절할 수 있습니다:
| 설정값 | 동작 방식 | 최적 사용 사례 |
|---|---|---|
0.3 |
키워드 70%, 시맨틱 30% | 식별자(ID), 코드, 부품 번호가 다수 포함된 문서 |
0.5 |
동일 가중치 | 일반적인 혼합 콘텐츠 |
0.7 |
시맨틱 70%, 키워드 30% | 서술형 문서, 정책, 절차서 |
1.0 |
순수 벡터 검색 | 정확한 용어가 중요하지 않은 개념 중심 Q&A |
0.0 |
순수 키워드 검색 | ID 또는 코드를 통한 정확한 조회 |
품질 문서(SOP, CAPA, 일탈 보고서 등)의 경우 0.7로 시작하여 실제 사용자의 검색 질의 패턴에 따라 조정하는 것을 권장합니다. 사용자가 문서 ID나 부품 번호로 자주 검색한다면 0.5로 낮추십시오.
서비스 접근 엔드포인트 (Access points)
| 서비스 | URL | 인증 정보 |
|---|---|---|
| RAG API | http://localhost:8000/docs |
없음 (Swagger UI) |
| RustFS 콘솔 | http://localhost:9001 |
admin / admin123456 |
| Ollama API | http://localhost:11434 |
없음 |
주의해야 할 함정 (What to avoid)
고정 크기 청킹(fixed-size chunking)을 지양하십시오. 500자마다 무조건 자르는 방식은 문장을 중간에 끊어버리고 본문과 제목을 분리시킵니다. 문단과 문장 경계를 존중하는 구조 인식형 청킹을 사용해야 합니다.
BM25 계층을 생략하지 마십시오. 벡터 전용 RAG는 단 한 글자만 다른 문서 ID, 부품 번호, 규제 조항 참조를 쉽게 혼동합니다. 규제 준수 문서에서 하이브리드 검색은 선택이 아닌 필수입니다.
초기 인덱싱에 클라우드 임베딩 모델을 사용하지 마십시오. 이 아키텍처의 핵심 가치는 데이터 주권(Data Sovereignty)입니다. 원본 문서 텍스트는 사내 네트워크 밖으로 절대 유출되지 않아야 합니다. 임베딩 모델은 반드시 자체 인프라의 Ollama에서 구동하십시오.
프로덕션 환경에서 프론티어 모델을 CPU로 실행하지 마십시오. 32B 파라미터 모델을 CPU로 추론하면 토큰 하나를 생성하는 데 수 초가 소요됩니다. 테스트용으로는 가능하지만, 실제 사용자가 상호작용하는 프로덕션 환경에서는 GPU를 장착하거나 자체 호스팅 API 엔드포인트를 사용해야 합니다.
핵심 요약 및 결론 (The bottom line)
본 아키텍처 스택은 4개의 Docker 서비스로 완벽히 컨테이너화되어, 최신 하이브리드 검색과 프론티어 모델급 텍스트 생성 품질을 갖춘 100% 로컬 임베딩 파이프라인을 제공합니다. 원본 문서는 사내 인프라를 결코 벗어나지 않습니다. 검색 프로세스는 투명하며 세밀하게 튜닝할 수 있습니다. 프론티어 모델은 단 하나의 환경 변수 변경만으로 손쉽게 교체 가능합니다.
규제 환경의 품질 문서 관리에서 하이브리드 RAG는 단순한 부가 기능이 아닙니다. 그것은 반드시 갖추어야 할 최소 실행 가능 검색 아키텍처(Minimum Viable Retrieval Architecture)입니다. 순수 벡터 검색은 SOP 번호를 놓치고, 순수 키워드 검색은 개념적 질의를 놓칩니다. 하이브리드 검색만이 두 마리 토끼를 모두 잡을 수 있습니다.
연구 노트: [[hybrid-rag-docker-tutorial-analysis]]
Saram Consulting