从日期和文本构建特征
从原始字段中提取有效信号
从日期和文本构建特征 是 CoddyKit 上的免费 Data Science Academy 课时。 这是第 4 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Data Science Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Data Science Academy 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
Raw Fields Hide Signal
A timestamp or a free-text note holds rich clues a model cannot read directly. Feature engineering pulls that signal into usable columns. 🔍
Parse the Date First
Before mining a date, convert the string with to_datetime. Only a real datetime exposes the handy parts you want to extract.
df['ts'] = pd.to_datetime(df['ts'])The dt Accessor
The dt accessor unlocks date parts on a datetime column. Reach for it to grab year, month, day, and more in one line.
df['year'] = df['ts'].dt.year
df['month'] = df['ts'].dt.monthDay of Week
Pull dayofweek to capture weekly rhythm, where Monday is 0 and Sunday is 6. Behavior often shifts a lot by weekday.
df['weekday'] = df['ts'].dt.dayofweekWeekend Flag
Turn the weekday into a simple boolean: is this a weekend? Such yes-or-no flags are easy for models to learn from.
df['is_weekend'] = df['ts'].dt.dayofweek >= 5Cyclical Time
Hour 23 sits right next to hour 0, yet the numbers look far apart. Cyclical encoding with sine and cosine fixes that wrap-around.
Text: Length and Counts
Start simple with text. The character length of a review or its word count is often a surprisingly strong feature.
df['len'] = df['review'].str.len()
df['words'] = df['review'].str.split().str.len()Contains a Keyword
Flag whether text holds a key term using str.contains. A column for words like refund can capture clear intent.
df['has_refund'] = df['review'].str.contains('refund')Bag of Words
To use words themselves, count them. A bag-of-words turns text into one column per term holding how often it appears.
from sklearn.feature_extraction.text import CountVectorizer
X = CountVectorizer().fit_transform(df['review'])TF-IDF Weighs Words
Plain counts overrate common words. TF-IDF rewards terms that are frequent in one row yet rare overall, so the distinctive words stand out.
from sklearn.feature_extraction.text import TfidfVectorizer
X = TfidfVectorizer().fit_transform(df['review'])Avoid Leaky Features
Never build a feature from information unavailable at prediction time. A leaky column inflates scores in testing, then fails in the real world.
Quick Check
You have a datetime column and want the weekday number. What do you use?
Recap: Date and Text Features
You mined raw fields into signal: parse dates then use dt for parts and flags, and turn text into lengths, keyword flags, or word counts. 🎉
常见问题解答
「从日期和文本构建特征」课时是免费的吗?
是的 — 「从日期和文本构建特征」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Data Science Academy 课程的其余内容,请升级到 CoddyKit PRO。 Data Science Academy 课程共包含 4 节课。
「从日期和文本构建特征」这节课中我会学到什么?
从原始字段中提取有效信号 你通过在浏览器中直接运行的动手代码来练习 Data Science Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 Data Science Academy 需要有经验吗?
无需任何先前经验。CoddyKit 上的 Data Science Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 4 节课,共 4 节。
「从日期和文本构建特征」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 Data Science Academy 课中编写并运行代码吗?
能。每节 Data Science Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。
此课程中的所有课时
- 将数字区间划分为类别
- 编码分类列
- 缩放和标准化数值
- 从日期和文本构建特征