Tokens, Spans und Doc-Objekte
Die Datenstrukturen von spaCy durchlaufen
Tokens, Spans und Doc-Objekte ist eine kostenlose NLP Academy-Lektion auf CoddyKit. Dies ist Lektion 3 von 4. Du kannst die komplette Lektion unten kostenlos lesen – dann übst du sie direkt im Browser mit einem integrierten Code-Editor und einem KI-Tutor rund um die Uhr. Sie ist Teil des NLP Academy-Lernpfads, und dein Fortschritt wird über Web und CoddyKit-App synchronisiert. Der NLP Academy-Kurs umfasst insgesamt 4 Lektionen.
Teile dieser Lektion wurden noch nicht übersetzt und werden auf Englisch angezeigt.
Three Core Objects
spaCy gives you three building blocks: the Doc, the Token, and the Span. Learn these and the whole library opens up.
The Doc Container
A Doc is the full processed document. It holds the original text plus every token and all the analysis spaCy produced.
doc = nlp("Maria leads the data team.")Index Like a List
You can grab one item by position because a Doc is indexable. Indexing returns a single Token, just like a Python list.
first = doc[0]What a Token Holds
A Token is one unit of text plus rich metadata: its part of speech, its lemma, and whether it is punctuation.
print(doc[0].text, doc[0].pos_)Useful Token Flags
Tokens carry handy flags like is_stop, is_alpha, and is_punct. They let you filter words without writing your own checks.
for t in doc:
print(t.text, t.is_stop)Slice to Get a Span
Slice a Doc and you get a Span, a contiguous run of tokens. It is perfect for phrases like a person or a place.
phrase = doc[0:2]Spans Keep Context
A Span is not a copy. It still points back into the Doc, so it knows its position and shares all the original analysis.
Entities Are Spans
Named entities show up as Spans in doc.ents. Each one bundles the words plus a label like PERSON or ORG.
for ent in doc.ents:
print(ent.text, ent.label_)Noun Chunks
spaCy also offers noun_chunks, base noun phrases served as Spans. They are a quick way to pull out the things a sentence talks about.
list(doc.noun_chunks)Sentences as Spans
Need sentences? Iterate doc.sents. Each sentence comes back as a Span, so segmentation is already handled for you.
for sent in doc.sents:
print(sent.text)Original Text Stays
Through all of this the original string is preserved in doc.text. Tokens and spans are views, never destructive edits.
Quick Check
What do you get when you slice a Doc with doc[0:3]?
Recap
The Doc holds everything, indexing gives one Token, and slicing gives a Span. Entities, noun chunks, and sentences all arrive as Spans. 🧩
Häufig gestellte Fragen
Ist die Lektion „Tokens, Spans und Doc-Objekte“ kostenlos?
Ja — der vollständige Text von „Tokens, Spans und Doc-Objekte“ ist hier im Web kostenlos zu lesen. Um sie interaktiv zu üben (integrierter Code-Editor und 24/7 KI-Tutor) und den Rest des NLP Academy-Kurses freizuschalten, upgrade auf CoddyKit PRO. Der NLP Academy-Kurs umfasst insgesamt 4 Lektionen.
Was lerne ich in „Tokens, Spans und Doc-Objekte“?
Die Datenstrukturen von spaCy durchlaufen Du übst NLP Academy mit praktischem Code, den du direkt im Browser ausführst, und ein 24/7 KI-Tutor beantwortet deine Fragen während du die Lektion bearbeitest.
Brauche ich Erfahrung, um NLP Academy zu starten?
Keine Vorkenntnisse erforderlich. NLP Academy auf CoddyKit ist für Anfänger bis fortgeschrittene Lernende strukturiert, sodass du hier starten oder von Anfang an beginnen und in deinem eigenen Tempo voranschreiten kannst. Dies ist Lektion 3 von 4.
Wie lange dauert die Lektion „Tokens, Spans und Doc-Objekte“?
Die meisten CoddyKit-Lektionen dauern etwa 5–10 Minuten. Jede ist kompakt und interaktiv, sodass du stetig Fortschritte machst und genau dort weitermachst, wo du aufgehört hast – im Web und in der App.
Kann ich in dieser NLP Academy-Lektion Code schreiben und ausführen?
Ja. Jede NLP Academy-Lektion enthält einen integrierten Code-Editor, sodass du echten Code direkt in deinem Browser schreibst und ausführst und sofort KI-Feedback erhältst — ohne lokale Einrichtung erforderlich.
Alle Lektionen in diesem Kurs
- Warum spaCy für reale Projekte geeignet ist
- Ein Modell laden und ein Doc verarbeiten
- Tokens, Spans und Doc-Objekte
- Die Pipeline anpassen