Token, Span e oggetti Doc
Esplorare le strutture dati di spaCy
Token, Span e oggetti Doc è una lezione NLP Academy gratuita su CoddyKit. Questa è la lezione 3 di 4. Puoi leggere la lezione completa qui gratuitamente — poi esercitati direttamente nel browser con un editor di codice integrato e un tutor IA disponibile 24/7. Fa parte del percorso di apprendimento NLP Academy, e i tuoi progressi si sincronizzano tra il web e l'app CoddyKit. Il corso NLP Academy include 4 lezioni in totale.
Parti di questa lezione non sono ancora state tradotte e vengono mostrate in inglese.
Three Core Objects
spaCy gives you three building blocks: the Doc, the Token, and the Span. Learn these and the whole library opens up.
The Doc Container
A Doc is the full processed document. It holds the original text plus every token and all the analysis spaCy produced.
doc = nlp("Maria leads the data team.")Index Like a List
You can grab one item by position because a Doc is indexable. Indexing returns a single Token, just like a Python list.
first = doc[0]What a Token Holds
A Token is one unit of text plus rich metadata: its part of speech, its lemma, and whether it is punctuation.
print(doc[0].text, doc[0].pos_)Useful Token Flags
Tokens carry handy flags like is_stop, is_alpha, and is_punct. They let you filter words without writing your own checks.
for t in doc:
print(t.text, t.is_stop)Slice to Get a Span
Slice a Doc and you get a Span, a contiguous run of tokens. It is perfect for phrases like a person or a place.
phrase = doc[0:2]Spans Keep Context
A Span is not a copy. It still points back into the Doc, so it knows its position and shares all the original analysis.
Entities Are Spans
Named entities show up as Spans in doc.ents. Each one bundles the words plus a label like PERSON or ORG.
for ent in doc.ents:
print(ent.text, ent.label_)Noun Chunks
spaCy also offers noun_chunks, base noun phrases served as Spans. They are a quick way to pull out the things a sentence talks about.
list(doc.noun_chunks)Sentences as Spans
Need sentences? Iterate doc.sents. Each sentence comes back as a Span, so segmentation is already handled for you.
for sent in doc.sents:
print(sent.text)Original Text Stays
Through all of this the original string is preserved in doc.text. Tokens and spans are views, never destructive edits.
Quick Check
What do you get when you slice a Doc with doc[0:3]?
Recap
The Doc holds everything, indexing gives one Token, and slicing gives a Span. Entities, noun chunks, and sentences all arrive as Spans. 🧩
Domande Frequenti
La lezione «Token, Span e oggetti Doc» è gratuita?
Sì — il testo completo di «Token, Span e oggetti Doc» è gratuito qui sul web. Per esercitarvi in modo interattivo (un editor di codice integrato e un tutor IA 24/7) e sbloccare il resto del corso NLP Academy, passa a CoddyKit PRO. Il corso NLP Academy include 4 lezioni in totale.
Cosa imparerò in «Token, Span e oggetti Doc»?
Esplorare le strutture dati di spaCy Eserciti NLP Academy con codice pratico che esegui direttamente nel browser, e un tutor IA 24/7 risponde alle tue domande mentre lavori sulla lezione.
Ho bisogno di esperienza per iniziare NLP Academy?
Non è richiesta alcuna esperienza precedente. NLP Academy su CoddyKit è strutturato per principianti e studenti avanzati, quindi puoi iniziare da qui o dall'inizio e procedere al tuo ritmo. Questa è la lezione 3 di 4.
Quanto tempo richiede la lezione «Token, Span e oggetti Doc»?
La maggior parte delle lezioni CoddyKit richiede circa 5–10 minuti. Ogni lezione è breve e interattiva, quindi fai progressi costanti e riprendi esattamente da dove hai lasciato su web e app.
Posso scrivere ed eseguire codice in questa lezione NLP Academy?
Sì. Ogni lezione NLP Academy include un editor di codice integrato, quindi scrivi ed esegui codice reale direttamente nel tuo browser e ricevi feedback istantaneo dall'IA — nessuna configurazione locale necessaria.
Tutte le lezioni di questo corso
- Perché spaCy è adatto ai progetti reali
- Caricare un modello ed elaborare un Doc
- Token, Span e oggetti Doc
- Personalizzare la pipeline