0Pricing
NLP Academy · บทเรียน

โทเคน ช่วงข้อความ และออบเจ็กต์เอกสาร

สำรวจโครงสร้างข้อมูลของ spaCy

โทเคน ช่วงข้อความ และออบเจ็กต์เอกสาร เป็นบทเรียน NLP Academy ฟรีบน CoddyKit นี่คือบทเรียนที่ 3 จากทั้งหมด 4 บทเรียน คุณสามารถอ่านบทเรียนทั้งหมดด้านล่างฟรี — จากนั้นลองปฏิบัติด้วยตัวคุณเองในเบราว์เซอร์พร้อมตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7 บทเรียนนี้เป็นส่วนหนึ่งของเส้นทางการเรียน NLP Academy และความก้าวหน้าของคุณจะซิงค์ข้ามเว็บและแอป CoddyKit คอร์ส NLP Academy มีบทเรียนทั้งหมด 4 บทเรียน

บางส่วนของบทเรียนนี้ยังไม่ได้รับการแปล และแสดงเป็นภาษาอังกฤษ

Three Core Objects

spaCy gives you three building blocks: the Doc, the Token, and the Span. Learn these and the whole library opens up.

The Doc Container

A Doc is the full processed document. It holds the original text plus every token and all the analysis spaCy produced.

doc = nlp("Maria leads the data team.")

Index Like a List

You can grab one item by position because a Doc is indexable. Indexing returns a single Token, just like a Python list.

first = doc[0]

What a Token Holds

A Token is one unit of text plus rich metadata: its part of speech, its lemma, and whether it is punctuation.

print(doc[0].text, doc[0].pos_)

Useful Token Flags

Tokens carry handy flags like is_stop, is_alpha, and is_punct. They let you filter words without writing your own checks.

for t in doc:
    print(t.text, t.is_stop)

Slice to Get a Span

Slice a Doc and you get a Span, a contiguous run of tokens. It is perfect for phrases like a person or a place.

phrase = doc[0:2]

Spans Keep Context

A Span is not a copy. It still points back into the Doc, so it knows its position and shares all the original analysis.

Entities Are Spans

Named entities show up as Spans in doc.ents. Each one bundles the words plus a label like PERSON or ORG.

for ent in doc.ents:
    print(ent.text, ent.label_)

Noun Chunks

spaCy also offers noun_chunks, base noun phrases served as Spans. They are a quick way to pull out the things a sentence talks about.

list(doc.noun_chunks)

Sentences as Spans

Need sentences? Iterate doc.sents. Each sentence comes back as a Span, so segmentation is already handled for you.

for sent in doc.sents:
    print(sent.text)

Original Text Stays

Through all of this the original string is preserved in doc.text. Tokens and spans are views, never destructive edits.

Quick Check

What do you get when you slice a Doc with doc[0:3]?

Recap

The Doc holds everything, indexing gives one Token, and slicing gives a Span. Entities, noun chunks, and sentences all arrive as Spans. 🧩

คำถามที่พบบ่อย

บทเรียน “โทเคน ช่วงข้อความ และออบเจ็กต์เอกสาร” ฟรีหรือไม่

ใช่ — ข้อความเต็มของ “โทเคน ช่วงข้อความ และออบเจ็กต์เอกสาร” ฟรีให้อ่านที่นี่บนเว็บ เพื่อปฏิบัติแบบโต้ตอบ (ตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7) และปลดล็อคส่วนที่เหลือของคอร์ส NLP Academy ให้อัปเกรดเป็น CoddyKit PRO คอร์ส NLP Academy มีบทเรียนทั้งหมด 4 บทเรียน

คุณจะเรียนรู้อะไรในบทเรียน “โทเคน ช่วงข้อความ และออบเจ็กต์เอกสาร”

สำรวจโครงสร้างข้อมูลของ spaCy คุณปฏิบัติ NLP Academy ด้วยโค้ดที่ใช้งานได้จริงที่คุณเรียกใช้โดยตรงในเบราว์เซอร์ และติวเตอร์ AI ตลอด 24/7 ตอบคำถามของคุณขณะที่คุณไปผ่านบทเรียน

คุณต้องมีประสบการณ์ก่อนที่จะเริ่มเรียน NLP Academy หรือไม่

ไม่จำเป็นต้องมีประสบการณ์มาก่อน NLP Academy บน CoddyKit ออกแบบมาสำหรับผู้เริ่มต้นไปจนถึงผู้เรียนขั้นสูง คุณสามารถเริ่มต้นที่นี่หรือเริ่มจากตัวแรกและเรียนด้วยความเร็วของคุณเอง นี่คือบทเรียนที่ 3 จากทั้งหมด 4 บทเรียน

บทเรียน “โทเคน ช่วงข้อความ และออบเจ็กต์เอกสาร” ใช้เวลานานแค่ไหน

บทเรียน CoddyKit ส่วนใหญ่ใช้เวลาประมาณ 5–10 นาที แต่ละบทเรียนจึงสั้นและเป็นแบบโต้ตอบ คุณสามารถก้าวหน้าอย่างต่อเนื่องและกลับมาเรียนต่อจากตรงที่เพิ่งหยุดบนเว็บและแอปได้เลย

ฉันเขียนและรันโค้ดในบทเรียน NLP Academy นี้ได้ไหม

ได้ บทเรียน NLP Academy ทุกบทมีตัวแก้ไขโค้ดในตัว คุณจึงเขียนและรันโค้ดจริงได้เลยในเบราว์เซอร์ และได้รับข้อเสนอแนะจาก AI ในทันที — ไม่ต้องติดตั้งในเครื่องของคุณ

บทเรียนทั้งหมดในหลักสูตรนี้

  1. เหตุใดจึงเลือก spaCy สำหรับโครงการจริง
  2. โหลดโมเดลและประมวลผลเอกสาร
  3. โทเคน ช่วงข้อความ และออบเจ็กต์เอกสาร
  4. ปรับแต่งไปป์ไลน์
← กลับไปที่ NLP Academy