Trucchi per la tokenizzazione con regex
Dividere il testo esattamente come desidera
Trucchi per la tokenizzazione con regex è una lezione NLP Academy gratuita su CoddyKit. Questa è la lezione 4 di 4. Puoi leggere la lezione completa qui gratuitamente — poi esercitati direttamente nel browser con un editor di codice integrato e un tutor IA disponibile 24/7. Fa parte del percorso di apprendimento NLP Academy, e i tuoi progressi si sincronizzano tra il web e l'app CoddyKit. Il corso NLP Academy include 4 lezioni in totale.
Parti di questa lezione non sono ancora state tradotte e vengono mostrate in inglese.
Tokenizing With Patterns
You can tokenize text by describing what a token looks like, then letting regex find every piece that fits your definition.
Match Words Directly
The pattern \w+ with findall grabs every run of word characters, giving you a clean list of words and ignoring the spaces between them.
re.findall("\w+", "Hello, world!")Split Instead of Match
The re.split function breaks text wherever a pattern appears. Splitting on whitespace turns a sentence into a list of rough tokens fast.
re.split("\s+", "one two three")Split on Multiple Separators
With a character class, re.split can break on many separators at once, like spaces, commas, and semicolons in a single sweep.
re.split("[ ,;]+", "a, b;c d")Keep Punctuation as Tokens
Sometimes punctuation matters. A pattern like \w+|[^\w\s] captures words and lone symbols separately, so nothing is silently dropped.
The Alternation Operator
The pipe means or. The pattern cat|dog matches either word, letting one regex describe several token shapes you care about.
re.findall("cat|dog", "a dog, a cat")Match Numbers and Decimals
Numbers need their own rule. The pattern \d+\.?\d* matches whole numbers and decimals, so prices and amounts stay together as one token.
re.findall("\d+\.?\d*", "buy 3 for 4.50")Handle Contractions
Naive splitting wrecks words like do not in its short form. A smarter token pattern can keep apostrophes inside a word where they belong.
Word Boundaries Help
The \b anchor marks a word boundary, the edge between a word and a non-word. It helps you grab whole words without grabbing neighbors.
Build a Token Pattern
A practical tokenizer often combines rules with alternation: match URLs, then numbers, then words, then symbols, in priority order.
Compile for Speed
If you reuse a pattern a lot, re.compile turns it into a reusable object. It reads cleaner and runs faster across many strings.
tok = re.compile("\w+")Quick Check
You want to break a string anywhere one or more spaces appear. Which function fits best?
Recap: Regex Tokenizers
You used findall, split, alternation, and compile to turn raw text into exactly the tokens you want. Regex gives you full control. 🎯
Domande Frequenti
La lezione «Trucchi per la tokenizzazione con regex» è gratuita?
Sì — il testo completo di «Trucchi per la tokenizzazione con regex» è gratuito qui sul web. Per esercitarvi in modo interattivo (un editor di codice integrato e un tutor IA 24/7) e sbloccare il resto del corso NLP Academy, passa a CoddyKit PRO. Il corso NLP Academy include 4 lezioni in totale.
Cosa imparerò in «Trucchi per la tokenizzazione con regex»?
Dividere il testo esattamente come desidera Eserciti NLP Academy con codice pratico che esegui direttamente nel browser, e un tutor IA 24/7 risponde alle tue domande mentre lavori sulla lezione.
Ho bisogno di esperienza per iniziare NLP Academy?
Non è richiesta alcuna esperienza precedente. NLP Academy su CoddyKit è strutturato per principianti e studenti avanzati, quindi puoi iniziare da qui o dall'inizio e procedere al tuo ritmo. Questa è la lezione 4 di 4.
Quanto tempo richiede la lezione «Trucchi per la tokenizzazione con regex»?
La maggior parte delle lezioni CoddyKit richiede circa 5–10 minuti. Ogni lezione è breve e interattiva, quindi fai progressi costanti e riprendi esattamente da dove hai lasciato su web e app.
Posso scrivere ed eseguire codice in questa lezione NLP Academy?
Sì. Ogni lezione NLP Academy include un editor di codice integrato, quindi scrivi ed esegui codice reale direttamente nel tuo browser e ricevi feedback istantaneo dall'IA — nessuna configurazione locale necessaria.
Tutte le lezioni di questo corso
- Le espressioni regolari in 5 minuti
- Trovare email e URL
- Gruppi di cattura e sostituzioni
- Trucchi per la tokenizzazione con regex