Trong bài viết trước, chúng ta đã xây dựng một đường ống phân tích văn bản sử dụng quanteda. Quy trình đó tạo ra ma trận đặc trưng văn bản (DFM), áp dụng trọng số TF-IDF và đo lường sự tương đồng qua cosine similarity. Tuy nhiên, phương pháp này bộc lộ một hạn chế căn bản: cosine similarity tính trên các vector TF-IDF thưa chỉ phản ánh sự trùng lặp từ vựng, chứ không nắm bắt được ý nghĩa ngữ nghĩa.
Xét ví dụ hai câu: "Người vay nợ không thực hiện nghĩa vụ trả nợ" và "Khách hàng không đáp ứng được cam kết trả nợ". Sau khi loại từ dừng, chúng không chia sẻ từ vựng nào, dẫn đến độ tương đồng TF-IDF gần như bằng 0. Dù vậy, bất kỳ аналитик nào cũng nhận ra chúng mang ý nghĩa như nhau.
Giải pháp nằm ở vector nhúng dày đặc (dense embeddings) — biểu diễn thần kinh nơi các văn bản mang ý nghĩa gần nhau được ánh xạ đến các vị trí lân cận trong không gian vectơ. Nhờ đó, chúng ta có thể đo lường sự tương đồng ngữ nghĩa ngay cả khi không có từ chung.
Bài viết này sẽ hướng dẫn xây dựng một đường ống Tạo sinh tăng cường truy xuất (RAG) trong R, sử dụng ragnar cho kho truy xuất và ellmer cho tương tác mô hình ngôn ngữ lớn (LLM), minh họa sức mạnh của nhúng dày đặc trong truy xuất và sinh câu trả lời. Toàn bộ quy trình chạy cục bộ nhờ Ollama.
Vấn đề của TF-IDF và tương đồng cosine
Để thấy rõ hạn chế, ta so sánh trực tiếp hai phương pháp trên cùng một tập câu.
1library(quanteda)
2library(quanteda.textstats)
3
4# Hai câu ngữ nghĩa giống hệt nhau nhưng từ vựng hoàn toàn khác
5sentences <- c(
6 doc1 = "The borrower defaulted on their loan obligations.",
7 doc2 = "The client failed to meet repayment commitments.",
8 doc3 = "The borrower defaulted on their loan obligations." # giống hệt doc1
9)
10
11corp <- corpus(sentences)
12
13sim_tfidf <- corp |>
14 tokens(remove_punct = TRUE) |>
15 tokens_tolower() |>
16 tokens_remove(stopwords("english")) |>
17 dfm() |>
18 dfm_tfidf() |>
19 textstat_simil(method = "cosine", margin = "documents")
20
21as.matrix(sim_tfidf)1 doc1 doc2 doc3
2doc1 1 0 1
3doc2 0 1 0
4doc3 1 0 1Cặp doc1 - doc3 đạt 1.0 (giống hệt). Cặp doc1 - doc2 bằng 0, dù ý nghĩa y nhau. Đó là trần trụi giới hạn của việc chỉ dựa vào trùng từ.
Nhúng dày đặc: Biểu diễn ý nghĩa thay vì từ vựng
Vector nhúng dày đặc thay thế mỗi văn bản bằng một vectơ số độ dài cố định do mạng nơ-ron tạo ra. Mô hình nhúng được huấn luyện để đặt các văn bản cùng ngữ nghĩa gần nhau trong không gian vectơ, bất kể từ vựng có trùng nhau hay không.
1# Nhúng cùng ba câu bằng mô hình Ollama cục bộ
2embed_fn <- embed_ollama(model = "nomic-embed-text")
3
4embeddings <- embed_fn(as.character(sentences))
5
6# Hàm tính cosine similarity
7cosine_sim <- function(a, b) sum(a * b) / (sqrt(sum(a^2)) * sqrt(sum(b^2)))
8
9sim_12 <- cosine_sim(embeddings[1, ], embeddings[2, ])
10sim_13 <- cosine_sim(embeddings[1, ], embeddings[3, ])
11
12tibble(
13 pair = c("doc1 vs doc2 (khác từ, cùng ý)",
14 "doc1 vs doc3 (giống hệt)"),
15 tfidf_cosine = c(0.000, 1.000),
16 dense_cosine = round(c(sim_12, sim_13), 3)
17)1# A tibble: 2 × 3
2 pair tfidf_cosine dense_cosine
3 <chr> <dbl> <dbl>
41 doc1 vs doc2 (khác từ, cùng ý) 0 0.733
52 doc1 vs doc3 (giống hệt) 1 1Kết quả: doc1 so với doc2 giờ đạt khoảng 0.73 — phản ánh đúng sự tương đồng ngữ nghĩa — trong khi doc1 vs doc3 vẫn gần 1.0.
Chuẩn bị bộ ngữ liệu và chia nhỏ (chunking)
Chúng ta sẽ dùng một tập văn bản tổng hợp trích dẫn từ chính sách tín dụng, khung IFRS 9 và Basel III — loại tài liệu thường thấy trong tài liệu mô hình nội bộ.
1credit_docs <- tibble(
2 id = 1:15,
3 source = c(
4 rep("IFRS 9", 5),
5 rep("Basel III", 5),
6 rep("Credit Policy", 5)
7 ),
8 text = c(
9 # Trích dẫn IFRS 9
10 "A financial asset is in Stage 1 if there has been no significant increase in credit risk since initial recognition.",
11 "Stage 2 classification is triggered when there is a significant increase in credit risk relative to initial recognition, even if no actual default has occurred.",
12 "Stage 3 assets are those for which objective evidence of impairment exists, including default events or bankruptcy proceedings.",
13 "The 12-month Expected Credit Loss is the portion of lifetime ECL resulting from default events possible within 12 months of the reporting date.",
14 "Lifetime ECL represents the expected credit losses resulting from all possible default events over the expected life of the financial instrument.",
15
16 # Trích dẫn Basel III
17 "The Capital Conservation Buffer requires banks to hold a minimum Common Equity Tier 1 ratio of 2.5% above the regulatory minimum.",
18 "Counterparty credit risk arises from the failure of a counterparty to meet its contractual payment obligations.",
19 "The Liquidity Coverage Ratio ensures banks hold sufficient high-quality liquid assets to survive a 30-day stress scenario.",
20 "Pillar 2 requirements allow supervisors to impose additional capital requirements above the Pillar 1 minimum based on institution-specific risk profiles.",
21 "The Net Stable Funding Ratio measures the proportion of long-term assets funded by stable sources of funding over a one-year horizon.",
22
23 # Chính sách tín dụng
24 "Loan applicants with a debt-to-income ratio exceeding 45% are automatically declined under the current credit policy.",
25 "A payment more than 90 days past due triggers a non-performing loan classification and requires full provisioning.",
26 "Collateral revaluation must occur at minimum annually for secured exposures exceeding the materiality threshold.",
27 "The credit scorecard is recalibrated quarterly to reflect shifts in the applicant population and macroeconomic conditions.",
28 "Early arrears management is initiated when a borrower misses two consecutive scheduled payments."
29 )
30)
31
32credit_docsVới tài liệu thực tế (PDF, báo cáo dài), ragnar cung cấp markdown_chunk() để chia theo tiêu đề. Ở đây mỗi hàng đã đủ nhỏ, ta coi mỗi hàng là một đoạn (chunk) và bổ sung siêu dữ liệu.
1# Thực tế với văn bản lớn:
2# doc_text |> ragnar::markdown_chunk(level = 2) # tách theo tiêu đề ##
3
4# Với bộ ngữ liệu này, thêm metadata cho từng chunk
5chunks <- credit_docs |>
6 mutate(
7 chunk_id = paste0("chunk_", id),
8 char_count = nchar(text)
9 )
10
11chunks |> select(chunk_id, source, char_count, text)Mỗi chunk ở đây tương đương một hàng trong DFM. Quy trình truy xuất sau này sẽ đo cosine similarity giữa truy vấn và các chunk để tìm ngữ cảnh liên quan.
Xây dựng kho vectơ với ragnar
Kho vectơ lưu trữ các chunk cùng nhúng vectơ của chúng, hỗ trợ truy xuất nhanh dựa trên tương đồng vectơ dày đặc (VSS) và tần suất từ thưa (BM25). ragnar dùng DuckDB kèm mở rộng VSS, cho phép tạo kho, nhúng và lập chỉ mục chỉ trong vài dòng lệnh.
1# Tạo kho trong bộ nhớ (dùng đường dẫn file để bền vững)
2store <- ragnar_store_create(
3 location = ":memory:",
4 embed = embed_ollama(model = "nomic-embed-text"),
5 version = 1
6)
7
8# Chèn tài liệu
9ragnar_store_insert(
10 store,
11 chunks |> mutate(origin = source, hash = chunk_id) |> select(origin, hash, text)
12)
13
14# Xây dựng chỉ mục vectơ
15ragnar_store_build_index(store)Trực quan hóa không gian nhúng
Trước khi truy xuất, trực quan hóa giúp hiểu cấu trúc không gian nhúng.

1# Lấy toàn bộ chunk đã lưu từ DuckDB
2all_embeddings <- DBI::dbReadTable(store@con, "chunks")
3
4# PCA giảm chiều xuống 2D
5emb_matrix <- all_embeddings$embedding
6pca_result <- prcomp(emb_matrix, center = TRUE, scale. = FALSE)
7
8tibble(
9 source = chunks$source,
10 label = stringr::str_trunc(chunks$text, 50),
11 PC1 = pca_result$x[, 1],
12 PC2 = pca_result$x[, 2]
13) |>
14 ggplot(aes(x = PC1, y = PC2, colour = source, label = label)) +
15 ggrepel::geom_text_repel(size = 2.5, max.overlaps = 8) +
16 geom_point(size = 3, alpha = 0.8) +
17 labs(
18 title = "Không gian nhúng: Bộ ngữ liệu Chính sách tín dụng",
19 subtitle = "Shoot PCA của nhúng nomic-embed-text",
20 colour = "Nguồn"
21 ) +
22 theme_minimal()Các tài liệu cùng nguồn gom cụm lại với nhau. Nhúng đã nắm bắt cấu trúc lĩnh vực (IFRS 9 tách biệt Basel III) mặc dù mô hình không được huấn luyện chuyên biệt trên văn bản tài chính.
Truy xuất: VSS, BM25 và Hybrid
ragnar hỗ trợ ba chế độ truy xuất:
- VSS: Tương đồng cosine trên vectơ dày đặc (ngữ nghĩa).
- BM25: Tần suất từ thưa (trùng từ khóa).
- Hybrid: Hợp nhất hạng ngược (Reciprocal Rank Fusion) kết hợp cả hai.
1query <- "What triggers a move to Stage 2 under IFRS 9?"
2
3# Kết hợp (VSS + BM25)
4results_hybrid <- ragnar_retrieve(store, query, top_k = 2)
5
6# Chỉ VSS
7results_vss <- ragnar_retrieve_vss(store, query, top_k = 2)
8
9# Chỉ BM25
10results_bm25 <- ragnar_retrieve_bm25(store, query, top_k = 2)
11
12bind_rows(
13 results_vss |> mutate(method = "VSS", rank = row_number()),
14 results_bm25 |> mutate(method = "BM25", rank = row_number()),
15 results_hybrid |> mutate(method = "Hybrid", rank = row_number())
16) |>
17 select(method, rank, origin, text) |>
18 arrange(method, rank)1# A tibble: 7 × 4
2 method rank origin text
3 <chr> <int> <chr> <chr>
41 BM25 1 IFRS 9 Stage 2 classification is triggered when there is …
52 BM25 2 Credit Policy A payment more than 90 days past due triggers a no…
63 Hybrid 1 IFRS 9 Stage 2 classification is triggered when there is …
74 Hybrid 2 IFRS 9 A financial asset is in Stage 1 if there has been …
85 Hybrid 3 Credit Policy A payment more than 90 days past due triggers a no…
96 VSS 1 IFRS 9 Stage 2 classification is triggered when there is …
107 VSS 2 IFRS 9 A financial asset is in Stage 1 if there has been …Hybrid tìm cả trùng từ khóa chính xác (BM25) và ngữ nghĩa gần (VSS). VSS tập trung ngữ nghĩa, BM25 tập trung từ khóa.
Sinh câu trả lời có căn cứ với ellmer
Khi truy xuất đã ổn, ta dùng ngữ cảnh truy về để sinh câu trả lời có căn cứ. ellmer cung cấp giao diện xây dựng ứng dụng LLM trong R. Hàm ragnar_register_tool_retrieve() đăng ký kho như một công cụ (tool) cho LLM, cho phép mô hình tự chủ động kéo ngữ cảnh trước khi trả lời.
1# Khởi tạo chat (Ollama cục bộ)
2chat <- chat_ollama(
3 model = "llama3.1",
4 system = "You are a credit risk analyst. Answer questions using only the
5 information retrieved from the document store. If the answer is
6 not in the retrieved context, say so explicitly."
7)
8
9# Đăng ký kho vectơ như một công cụ truy xuất
10ragnar_register_tool_retrieve(chat, store)So sánh RAG và không RAG
Thử truy vấn cụ thể để thấy sự khác biệt.
1query <- "What triggers a loan to move to Stage 2 under IFRS 9?"
2
3# Có RAG (đã đăng ký công cụ truy xuất)
4answer_rag <- chat$chat(query)1The information retrieved does not specifically provide what triggers a loan to
2move to Stage 2 under IFRS 9. However, it suggests that the criteria for moving
3to Stage 2 might be closely related to an increase in default likelihood, as
4stated in one of the search results referencing "a significant increase in
5credit risk relative to initial recognition".
6
7Based solely on this result we cannot definitively state what triggers a loan
8under Stage 2 yet.1# Không RAG (chat mới, không công cụ, không ngữ cảnh)
2chat_naive <- chat_ollama(model = "llama3.1")
3answer_naive <- chat_naive$chat(query)1Under International Financial Reporting Standards (IFRS) 9, a loan is
2classified as an impaired asset and moved to Stage 2 when it experiences a
3significant increase in credit risk since the initial recognition of the loan.
4
5To be specific, a loan moves to Stage 2 if there are indicators that suggest a
6significant deterioration in the debtor's ability to repay the loan, which is
7considered to be more than insignificant. This can be due to various factors
8such as:
9
10* A decline in payments or any other sign that the borrower may default
11* A material increase in off-balance-sheet credit risk, such as a guarantee
12from a high-risk counterparty
13* A significant fall in credit ratings of the debtor
14
15Stage 2 is also known as "probable losses" (under IFRS 9) or "substandard
16loans" under many banking regulator's rules.1tibble(
2 approach = c("RAG (có căn cứ)", "Không RAG (nguyên sinh)"),
3 answer = c(answer_rag, answer_naive)
4)Câu trả lời RAG trích dẫn chính xác điều kiện kích hoạt Stage 2 từ ngữ liệu. Câu trả lời nguyên sinh tuy đúng общий nhưng thiếu căn cứ, và trong các trường hợp biên có thể ảo giác, dẫn ra điều kiện không tồn tại trong tài liệu nội bộ. Neo LLM vào ngữ cảnh truy xuất giúp giữ nó trung thực với nguồn.
Ứng dụng thực tế: Hệ thống hỏi đáp tài liệu mô hình
Một ứng dụng thực tiễn là xây dựng công cụ hỏi đáp nội bộ truy vấn tài liệu mô hình, khung quy định, sổ tay chính sách. Đường ống RAG sẽ truy xuất ngữ cảnh và sinh câu trả lời ngắn gọn.
1questions <- c(
2 "What triggers an IFRS 9 Stage 2 migration?",
3 "When is a loan classified as non-performing?",
4 "What is the Capital Conservation Buffer requirement?",
5 "How often is the credit scorecard recalibrated?"
6)
7
8# Hàm hỗ trợ: chạy truy vấn có căn cứ và trả về chuỗi
9ask_rag <- function(question) {
10 ch <- chat_ollama(
11 model = "llama3.1",
12 system = "Answer using only retrieved context. Be concise."
13 )
14 ragnar_register_tool_retrieve(ch, store)
15 ch$chat(question)
16}
17
18# Hỏi đáp hàng loạt dưới dạng tibble gọn gàng
19qa_results <- tibble(
20 question = questions,
21 answer = purrr::map_chr(questions, ask_rag)
22)1A significant increase in credit risk relative to initial recognition triggers
2an IFRS 9 Stage 2 migration.1According to the information retrieved, a loan is classified as non-performing
2when it has a payment that is more than 90 days past due.1The Capital Conservation Buffer requirement is to hold a minimum Common Equity
2Tier 1 ratio of 2.5% above the regulatory minimum, as per Basel III
3regulations. This means that banks must maintain a strong capital base to
4mitigate risks and ensure their financial stability.1The credit scorecard is recalibrated quarterly.1qa_results1# A tibble: 4 × 2
2 question answer
3 <chr> <chr>
41 What triggers an IFRS 9 Stage 2 migration? A significant increase i…
52 When is a loan classified as non-performing? According to the informa…
63 What is the Capital Conservation Buffer requirement? The Capital Conservation…
74 How often is the credit scorecard recalibrated? The credit scorecard is …Quy trình này lặp lại được mỗi khi tài liệu cập nhật, đảm bảo công cụ luôn mới. Thiết lập Ollama cục bộ nghĩa là không cần khóa API (API key) hay chi phí, lý tưởng cho dùng nội bộ.
✨ Giá trị đắt giá
- Nhúng dày đặc vượt qua giới hạn từ vựng của TF-IDF, nắm bắt ngữ nghĩa thực sự.
- RAG cục bộ với
ragnar+ellmer+Ollamacho phép triển khai hệ thống hỏi đáp bảo mật, không chi phí, dễ bảo trì. - Hybrid retrieval (VSS + BM25) cân bằng giữa tìm kiếm ngữ nghĩa và trùng từ khóa chính xác.
- Công cụ (tool) truy xuất cho phép LLM tự quyết định khi nào cần tra tài liệu, thay vì nhồi nhét toàn bộ ngữ cảnh vào prompt.
Câu hỏi tư duy: Nếu bộ ngữ liệu mở rộng lên hàng nghìn trang PDF quét (scan) chưa có lớp văn bản, bạn sẽ thiết kế đường ống tiền xử lý như thế nào trước khi đưa vào ragnar_store_insert()? Hãy thử triển khai bước OCR và làm sạch văn bản cho một tệp PDF mẫu.


