용어 빈도와 역문서 빈도
점수를 구성하는 두 부분을 알아봅니다
용어 빈도와 역문서 빈도은(는) CoddyKit의 무료 NLP Academy 강의입니다. 이것은 4개 중 2번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 NLP Academy 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. NLP Academy 강의에는 총 4개의 강의가 포함되어 있습니다.
이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.
Two Halves of One Score
TF-IDF is built from two pieces that multiply together: term frequency and inverse document frequency. Each half fixes a different weakness of raw counts.
Term Frequency, Plainly
Term frequency measures how often a word appears inside one document. More mentions of a word suggest that document leans toward that topic.
Normalizing TF
We often divide a word's count by the document length. This normalization stops long documents from looking important just because they have more words.
count = 3
doc_length = 50
tf = count / doc_length
print(round(tf, 3))TF Alone Isn't Enough
By itself, term frequency still rewards filler words that appear a lot. We need a second factor to punish words that show up everywhere.
Document Frequency
Document frequency counts how many documents contain a word at all. A high document frequency means the word is common across your whole collection.
Flip It: Inverse
We want rare words to score high, so we invert that count. Inverse document frequency rises when a word appears in few documents and falls when it is everywhere.
The Log Tames It
IDF uses a logarithm so the values do not explode for very rare words. The log keeps the scale smooth and comparable across terms.
import math
total_docs = 1000
docs_with_word = 10
idf = math.log(total_docs / docs_with_word)
print(round(idf, 3))Multiply Them Together
The final score is simply TF times IDF. A word wins only if it is frequent here and rare elsewhere, which is exactly the combination we wanted.
tf = 0.06
idf = 4.6
tfidf = tf * idf
print(round(tfidf, 3))What Scores High
A topic word like neuroscience in one article gets a high score. A word like the gets crushed because its IDF is nearly zero.
What Scores Low
Words appearing in almost every document earn tiny weights. TF-IDF automatically demotes this background noise without any manual stopword list.
Why It Works So Well
TF-IDF balances local importance against global rarity in one clean formula. That balance is why this weighting stayed a search and text staple for decades. ⭐
Quick Check
What does the inverse document frequency part actually reward?
Recap
TF measures local frequency, IDF rewards global rarity, and their product is TF-IDF. Together they spotlight words that truly define a document. ✅
자주 묻는 질문
“용어 빈도와 역문서 빈도” 강의는 무료인가요?
네 — “용어 빈도와 역문서 빈도” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 NLP Academy 강의 전체를 잠금 해제할 수 있습니다. NLP Academy 강의에는 총 4개의 강의가 포함되어 있습니다.
“용어 빈도와 역문서 빈도”에서 뭘 배우나요?
점수를 구성하는 두 부분을 알아봅니다 브라우저에서 직접 실행하는 실습 코드로 NLP Academy을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.
NLP Academy을(를) 시작하는 데 경험이 필요한가요?
사전 경험은 필요하지 않습니다. CoddyKit의 NLP Academy은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 2번째 강의입니다.
“용어 빈도와 역문서 빈도” 강의는 얼마나 걸리나요?
대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.
이 NLP Academy 강의에서 코드를 작성하고 실행할 수 있나요?
네. 모든 NLP Academy 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.
이 강의의 모든 강의
- 원시 카운트의 문제
- 용어 빈도와 역문서 빈도
- scikit-learn으로 TF-IDF 사용하기
- 가장 중요한 단어 찾기