TF-IDF और पद-आवृत्ति विश्लेषण
bind_tf_idf() की सहायता से दस्तावेज़ों में पदों के महत्त्व का स्कोर निकालें।
TF-IDF और पद-आवृत्ति विश्लेषण, CoddyKit पर R Academy का एक निःशुल्क पाठ है। यह 4 में से 2वाँ पाठ है। इस अध्ययन पथ के 3 तक कोई भी पाठ पूरा पढ़ना निःशुल्क है — इसके बाद CoddyKit PRO हर पाठ अनलॉक करता है, साथ ही अंतर्निर्मित कोड संपादक और चौबीसों घंटे एआई शिक्षक के साथ व्यावहारिक अभ्यास भी उपलब्ध कराता है। यह R Academy सीखने के मार्ग का हिस्सा है और आपकी प्रगति वेब तथा CoddyKit ऐप पर सिंक होती रहती है। R Academy पाठ्यक्रम में कुल 4 पाठ शामिल हैं।
TF-IDF क्या है
TF-IDF (Term Frequency–Inverse Document Frequency) मापता है कि किसी कॉर्पस के भीतर कोई शब्द किसी विशेष दस्तावेज़ के लिए कितना विशिष्ट है। TF-IDF का उच्च अंक बताता है कि शब्द उस दस्तावेज़ में बार-बार आता है, लेकिन सभी दस्तावेज़ों में दुर्लभ है—इसलिए यह दस्तावेज़ के प्रमुख विषयों का अच्छा संकेतक है।
# TF-IDF formula
# TF(w, d) = count of w in document d / total words in d
# IDF(w) = log(N / number of docs containing w)
# TF-IDF(w, d) = TF(w, d) * IDF(w)
# Manual example
N <- 4 # total documents
docs_with_word <- 1 # 'neural' appears in only 1 doc
tf <- 5 / 100 # word appears 5 times in 100-word doc
idf <- log(N / docs_with_word)
tfidf <- tf * idf
cat('TF: ', round(tf, 4), '\n')
cat('IDF: ', round(idf, 4), '\n')
cat('TF-IDF: ', round(tfidf, 4), '\n')टोकन गणना तैयार करना
TF-IDF निकालने से पहले आपके पास इन कॉलमों वाला डेटा फ़्रेम होना चाहिए: document (दस्तावेज़ पहचानकर्ता), word (टोकन) और n (गणना)। unnest_tokens() के बाद count(document, word) का उपयोग करें।
library(tidytext)
library(dplyr)
# Simulate a 4-document corpus on tech topics
docs <- tibble::tibble(
document = c(rep('ML', 3), rep('Stats', 3), rep('DB', 3), rep('Web', 3)),
text = c(
'machine learning neural networks training',
'deep learning models gradient descent',
'reinforcement learning reward policy agent',
'probability distributions hypothesis testing',
'bayesian inference regression analysis variance',
'statistics sampling confidence intervals normal',
'database sql joins indexes queries transactions',
'relational tables primary keys foreign constraints',
'nosql document store mongodb redis cassandra',
'html css javascript frontend react components',
'dom events browser api fetch requests',
'responsive design flexbox grid layout'
)
)
word_counts <- docs |>
unnest_tokens(word, text) |>
count(document, word, sort = TRUE)
cat('Document-word pairs:', nrow(word_counts), '\n')
print(head(word_counts, 8))bind_tf_idf: TF-IDF निकालना
bind_tf_idf(word, document, n) एक tidy गणना डेटा फ़्रेम लेता है और तीन कॉलम जोड़ता है: tf (शब्द आवृत्ति), idf (प्रतिलोम दस्तावेज़ आवृत्ति) और tf_idf (दोनों का गुणनफल)। यह डेटा फ़्रेम में मौजूद दस्तावेज़ों से IDF निकालता है।
library(tidytext)
library(dplyr)
docs <- tibble::tibble(
document = c(rep('ML', 2), rep('Stats', 2), rep('DB', 2)),
text = c(
'machine learning neural gradient',
'deep networks training epochs',
'probability distributions variance covariance',
'bayesian inference prior posterior',
'sql database tables joins indexes',
'queries transactions foreign primary'
)
)
tfidf <- docs |>
unnest_tokens(word, text) |>
count(document, word) |>
bind_tf_idf(word, document, n)
cat('Columns added:', names(tfidf), '\n')
print(head(arrange(tfidf, desc(tf_idf)), 8))arrange(desc(tf_idf)): शीर्ष शब्द
प्रत्येक दस्तावेज़ को सबसे अच्छी तरह दर्शाने वाले शब्द देखने के लिए घटते हुए tf_idf के आधार पर क्रमबद्ध करें। शून्य IDF वाले शब्द (जो हर दस्तावेज़ में मौजूद हैं) अपनी आवृत्ति चाहे जो हो, उनका TF-IDF शून्य होगा—वे दस्तावेज़ों में अंतर नहीं बता सकते।
library(tidytext)
library(dplyr)
docs <- tibble::tibble(
document = c(rep('Python', 3), rep('R', 3), rep('Julia', 3)),
text = c(
'python pandas numpy scipy programming',
'machine learning scikit tensorflow keras',
'jupyter notebooks scripts debugging modules',
'ggplot2 dplyr tidyverse statistics vectors',
'shiny rmarkdown knitr cran bioconductor',
'linear models factors dataframes packages',
'julia multiple dispatch type system macros',
'package ecosystem flux diffeq parallel',
'scientific computing performance benchmarks'
)
)
top_terms <- docs |>
unnest_tokens(word, text) |>
count(document, word) |>
bind_tf_idf(word, document, n) |>
arrange(desc(tf_idf))
cat('Top distinctive terms per language:\n')
print(head(top_terms[, c('document', 'word', 'tf_idf')], 9))प्रति-दस्तावेज़ शीर्ष TF-IDF शब्द
प्रत्येक दस्तावेज़ के सबसे विशिष्ट शीर्ष N शब्द निकालने के लिए group_by(document) |> slice_max(tf_idf, n = 5) का उपयोग करें। यह कॉर्पस का प्रभावी सारांश है—हर दस्तावेज़ की विशिष्ट शब्दावली स्पष्ट रूप से दिखाई देती है।
library(tidytext)
library(dplyr)
docs <- tibble::tibble(
document = c(rep('Cooking', 3), rep('Finance', 3), rep('Sports', 3)),
text = c(
'recipe ingredients oven bake flour butter',
'cooking temperature boil simmer saute',
'kitchen knife chopping herbs spices garnish',
'investment portfolio returns dividends stocks bonds',
'market equity risk hedge fund derivatives',
'inflation interest rate treasury balance sheet',
'goal tackle pass dribble goalkeeper penalty',
'tournament league championship trophy season',
'athlete training stamina sprint endurance'
)
)
top5 <- docs |>
unnest_tokens(word, text) |>
count(document, word) |>
bind_tf_idf(word, document, n) |>
group_by(document) |>
slice_max(tf_idf, n = 3) |>
ungroup()
print(top5[, c('document', 'word', 'tf_idf')])प्रति-दस्तावेज़ TF-IDF का दृश्यांकन
दस्तावेज़ के अनुसार बनाए गए TF-IDF अंकों के बार चार्ट से प्रत्येक दस्तावेज़ की विशिष्ट शब्दावली एक नज़र में दिखाई देती है। facet_wrap(~document, scales = 'free_y') का उपयोग करें ताकि प्रत्येक पैनल अपने शीर्ष शब्दों को स्वतंत्र रूप से दिखाए।
library(tidytext)
library(dplyr)
library(ggplot2)
docs <- tibble::tibble(
document = c(rep('AI', 3), rep('DB', 3), rep('Web', 3)),
text = c(
'neural networks gradient backpropagation',
'deep learning transformer attention bert',
'reinforcement policy reward exploration',
'sql database indexes joins transactions',
'relational schema normalization queries',
'nosql mongodb redis cassandra document',
'html css javascript react dom',
'frontend api fetch async components',
'responsive flexbox grid webpack bundle'
)
)
plot_data <- docs |>
unnest_tokens(word, text) |>
count(document, word) |>
bind_tf_idf(word, document, n) |>
group_by(document) |>
slice_max(tf_idf, n = 4) |>
ungroup()
ggplot(plot_data, aes(reorder(word, tf_idf), tf_idf, fill = document)) +
geom_col(show.legend = FALSE) +
facet_wrap(~document, scales = 'free_y') +
coord_flip() +
labs(x = NULL, y = 'TF-IDF', title = 'Top TF-IDF Terms per Topic') +
theme_minimal()IDF: सभी दस्तावेज़ों में आने वाले शब्द
हर दस्तावेज़ में आने वाले शब्दों के लिए IDF = log(N/N) = 0 होता है और इसलिए TF-IDF = 0 होता है। इससे विराम शब्दों की सूची की आवश्यकता के बिना पूरे कॉर्पस में सामान्य शब्दों का भार अपने-आप कम हो जाता है—हालाँकि बहुत छोटे कॉर्पस के लिए विराम शब्द हटाना अभी भी उपयोगी है।
library(tidytext)
library(dplyr)
docs <- tibble::tibble(
document = c('A', 'B', 'C'),
text = c(
'data analysis statistics regression probability',
'data engineering pipelines etl transformation',
'data visualisation ggplot2 charts dashboards'
)
)
tfidf <- docs |>
unnest_tokens(word, text) |>
count(document, word) |>
bind_tf_idf(word, document, n)
# 'data' appears in all 3 docs: tf_idf should be 0
data_rows <- filter(tfidf, word == 'data')
cat('TF-IDF for "data" across documents:\n')
print(data_rows[, c('document', 'word', 'idf', 'tf_idf')])दो दस्तावेज़ों की तुलना
TF-IDF दस्तावेज़ों की तुलना आसान बनाता है: किसी एक दस्तावेज़ में उच्च TF-IDF लेकिन दूसरे में कम या शून्य TF-IDF वाले शब्द खोजें, ताकि समझ सकें कि प्रत्येक दस्तावेज़ को क्या विशिष्ट बनाता है। TF-IDF तालिकाओं को word पर इनर जॉइन करें और अंकों की तुलना करें।
library(tidytext)
library(dplyr)
docs <- tibble::tibble(
document = c(rep('Doc1', 3), rep('Doc2', 3)),
text = c(
'bayesian statistics prior posterior mcmc',
'markov chain monte carlo sampling',
'probability distributions conjugate inference',
'convolutional neural network image classification',
'pooling activation relu softmax batch',
'convolution filter feature map stride padding'
)
)
tfidf <- docs |>
unnest_tokens(word, text) |>
count(document, word) |>
bind_tf_idf(word, document, n)
# Top 3 unique terms per document
tfidf |>
group_by(document) |>
slice_max(tf_idf, n = 3) |>
select(document, word, tf_idf) |>
print()जेन ऑस्टिन के उपन्यासों के साथ TF-IDF
janeaustenr पैकेज ऑस्टिन के छह उपन्यासों का पूरा टेक्स्ट उपलब्ध कराता है। यह TF-IDF का एक पारंपरिक मानक उदाहरण है: ऑस्टिन के प्रशंसक प्रत्येक उपन्यास की विशिष्ट शब्दावली पहचान सकते हैं—"wentworth" शब्द लगभग केवल Persuasion में मिलता है।
library(tidytext)
library(dplyr)
# Requires janeaustenr package
# library(janeaustenr)
# austen_books() returns: book, text
# Simulated mini-Austen corpus
mini_austen <- tibble::tibble(
book = c(rep('Sense', 3), rep('Pride', 3), rep('Emma', 3)),
text = c(
'elinor marianne dashwood willoughby colonel',
'edward ferrars barton cottage sister sense',
'brandon feelings attachment sensibility heart',
'darcy bennet bingley netherfield wickham',
'jane lizzy lydia longbourn pemberley',
'pride prejudice proposal marriage happiness',
'emma woodhouse knightley harriet weston',
'highbury match social governess niece',
'frank jane fairfax box hill picnic'
)
)
top <- mini_austen |>
unnest_tokens(word, text) |>
count(book, word) |>
bind_tf_idf(word, book, n) |>
group_by(book) |>
slice_max(tf_idf, n = 3) |>
select(book, word, tf_idf)
print(top)फ़ीचर इंजीनियरिंग के लिए TF-IDF
TF-IDF अंकों का उपयोग मशीन लर्निंग वर्गीकारकों में फ़ीचर के रूप में किया जा सकता है। tidy TF-IDF परिणाम को cast_dtm() या cast_sparse() का उपयोग करके दस्तावेज़-शब्द मैट्रिक्स में बदलें और उसे glmnet, xgboost या किसी अन्य वर्गीकारक को दें।
library(tidytext)
library(dplyr)
docs <- tibble::tibble(
document = c(rep('Positive', 4), rep('Negative', 4)),
text = c(
'excellent product love recommend quality',
'amazing fast delivery five star perfect',
'great value money satisfied happy return',
'wonderful experience purchase outstanding best',
'terrible waste money broken arrived damaged',
'horrible quality slow delivery disappointed awful',
'poor value cheap plastic flimsy broken',
'worst experience refund needed useless junk'
)
)
# TF-IDF as features
tfidf_wide <- docs |>
unnest_tokens(word, text) |>
count(document, word) |>
bind_tf_idf(word, document, n) |>
cast_dtm(document, word, tf_idf)
cat('DTM dimensions:', dim(tfidf_wide), '\n')
cat('(rows=documents, cols=unique words)\n')सीमाएँ और विकल्प
TF-IDF की कुछ सीमाएँ हैं: यह शब्दों को स्वतंत्र मानता है, शब्द-क्रम और संदर्भ को नज़रअंदाज़ करता है और दुर्लभ वर्तनी-त्रुटियों से प्रभावित हो सकता है। आधुनिक विकल्पों में शब्द एम्बेडिंग (word2vec, GloVe) और संदर्भात्मक एम्बेडिंग (BERT) शामिल हैं, लेकिन दस्तावेज़ खोज और वर्गीकरण के आधारभूत मॉडलों के लिए TF-IDF अभी भी उत्कृष्ट है।
# TF-IDF strengths and weaknesses summary
criteria <- data.frame(
Criterion = c('Speed', 'Interpretability', 'Context', 'Word order',
'Rare words', 'Scalability'),
TF_IDF = c('Fast', 'High', 'None', 'Ignored',
'High IDF (inflated)', 'Excellent'),
BERT = c('Slow', 'Low', 'Rich', 'Captured',
'Handled', 'Moderate')
)
print(criteria, row.names = FALSE)त्वरित जाँच
आपके कॉर्पस के सभी दस्तावेज़ों में कोई शब्द मौजूद है। उसका TF-IDF अंक क्या होगा?
पुनरावलोकन: TF-IDF विश्लेषण
मुख्य बातें:
- TF-IDF = शब्द आवृत्ति × प्रतिलोम दस्तावेज़ आवृत्ति; उच्च अंक = विशिष्ट शब्द
- कार्यप्रवाह:
unnest_tokens()→count(document, word)→bind_tf_idf(word, document, n) bind_tf_idf()tf,idfऔरtf_idfकॉलम जोड़ता हैarrange(desc(tf_idf))/slice_max(tf_idf, n = 5)प्रत्येक दस्तावेज़ के शीर्ष शब्द दिखाते हैं- सभी दस्तावेज़ों में मौजूद शब्दों का IDF = 0 और TF-IDF = 0 अपने-आप होता है
cast_dtm(document, word, tf_idf)ML के लिए दस्तावेज़-शब्द मैट्रिक्स में बदलता है- फ़ैसेट वाले बार चार्ट TF-IDF परिणामों का मानक दृश्यांकन हैं
library(tidytext)
library(dplyr)
# Minimal TF-IDF pipeline
tibble::tibble(
doc = c('A', 'A', 'B', 'B'),
text = c('neural networks deep', 'learning training', 'sql database query', 'joins indexes')
) |>
unnest_tokens(word, text) |>
count(doc, word) |>
bind_tf_idf(word, doc, n) |>
arrange(desc(tf_idf)) |>
print()एआई शिक्षक के साथ R सीखें — निःशुल्क
अपने ब्राउज़र में वास्तविक कोड लिखें और चलाएँ, चौबीसों घंटे एआई शिक्षक से तुरंत सहायता पाएँ, और वेब या ऐप पर वहीं से शुरू करें जहाँ आपने छोड़ा था।
- पाठ्यक्रम
- 43
- पाठ
- 159
अक्सर पूछे जाने वाले प्रश्न
क्या “TF-IDF और पद-आवृत्ति विश्लेषण” पाठ निःशुल्क है?
हाँ — R Academy अध्ययन पथ के 3 तक कोई भी पाठ, जिसमें “TF-IDF और पद-आवृत्ति विश्लेषण” भी शामिल है, यहाँ वेब पर पूरा पढ़ना निःशुल्क है। इसके बाद CoddyKit PRO हर पाठ अनलॉक करता है, साथ ही अंतर्निर्मित कोड संपादक और चौबीसों घंटे एआई शिक्षक के साथ इंटरैक्टिव अभ्यास भी उपलब्ध कराता है। R Academy पाठ्यक्रम में कुल 4 पाठ शामिल हैं।
“TF-IDF और पद-आवृत्ति विश्लेषण” में मैं क्या सीखूँगा?
bind_tf_idf() की सहायता से दस्तावेज़ों में पदों के महत्त्व का स्कोर निकालें। आप ब्राउज़र में सीधे चलाए जाने वाले व्यावहारिक कोड के साथ R Academy का अभ्यास करते हैं, और पाठ पूरा करते समय 24/7 एआई ट्यूटर आपके प्रश्नों के उत्तर देता है।
क्या R Academy शुरू करने के लिए मुझे किसी अनुभव की आवश्यकता है?
पहले के अनुभव की आवश्यकता नहीं है। CoddyKit पर R Academy शुरुआती से लेकर उन्नत शिक्षार्थियों तक सभी के लिए व्यवस्थित किया गया है, इसलिए आप यहीं से या शुरुआत से सीखना शुरू कर सकते हैं और अपनी गति से आगे बढ़ सकते हैं। यह 4 में से 2वाँ पाठ है।
“TF-IDF और पद-आवृत्ति विश्लेषण” पाठ पूरा करने में कितना समय लगता है?
CoddyKit का अधिकांश पाठ लगभग 5–10 मिनट में पूरा हो जाता है। हर पाठ छोटा और संवादात्मक है, इसलिए आप लगातार प्रगति करते हैं और वेब या ऐप पर वहीं से सीखना जारी रख सकते हैं जहाँ आपने छोड़ा था।
क्या मैं इस R Academy पाठ में कोड लिख और चला सकता हूँ?
हाँ। हर R Academy पाठ में एक अंतर्निर्मित कोड संपादक शामिल है, जिससे आप सीधे अपने ब्राउज़र में वास्तविक कोड लिख और चला सकते हैं और तुरंत एआई प्रतिक्रिया पा सकते हैं—स्थानीय सेटअप की आवश्यकता नहीं है।
इस पाठ्यक्रम के सभी पाठ
- टोकनीकरण और स्टॉप शब्द हटाना
- TF-IDF और पद-आवृत्ति विश्लेषण
- R में भावविश्लेषण
- LDA से विषय मॉडलिंग