Pipelines GDS et apprentissage automatique
Découvrez comment créer et gérer des pipelines de science des données appliquée aux graphes dans GDS, en les intégrant à des processus d’apprentissage automatique.
Pipelines GDS et apprentissage automatique est une leçon Neo4j Graph Database Fundamentals gratuite sur CoddyKit. Ceci est la leçon 3 sur 4. Tu peux lire la leçon complète ci-dessous gratuitement — puis la pratiquer en direct dans le navigateur avec un éditeur de code intégré et un tuteur IA 24/7. Elle fait partie du parcours d'apprentissage Neo4j Graph Database Fundamentals, et ta progression se synchronise sur le web et l'application CoddyKit. Le cours Neo4j Graph Database Fundamentals comprend 4 leçons au total.
Certaines parties de cette leçon n'ont pas encore été traduites et s'affichent en anglais.
What are GDS Pipelines?
Welcome to GDS Pipelines! In the Neo4j Graph Data Science (GDS) library, a pipeline is a structured workflow for common graph data science tasks.
Think of it as a blueprint that defines a sequence of steps, from feature engineering using graph algorithms to training and deploying machine learning models.
Why Use GDS Pipelines?
GDS Pipelines offer several key benefits:
- Reproducibility: Define your entire workflow once and reuse it.
- Automation: Streamline complex tasks involving multiple graph algorithms and ML steps.
- Operationalization: Easily deploy graph-based machine learning models for continuous prediction.
They help bridge the gap between experimentation and production.
Creating a Prediction Pipeline
One common pipeline type is the Node Property Prediction Pipeline. This helps predict a specific property on nodes based on other node features and graph structure.
Let's create an empty pipeline named 'myChurnPrediction':
CALL gds.pipeline.nodePropertyPrediction.create('myChurnPrediction')
YIELD pipeline
RETURN pipeline.name AS name, pipeline.type AS typeProjecting Your Graph Data
Before using a pipeline, you need to project your graph into GDS memory. This makes nodes and relationships accessible for algorithms and features.
Here, we project a simple 'User' graph with 'KNOWS' relationships:
CALL gds.graph.project(
'mySocialGraph',
'User',
{ KNOWS: { orientation: 'UNDIRECTED' } }
)
YIELD graphName, nodeCount, relationshipCount
RETURN graphName, nodeCount, relationshipCountAdding Node Property Features
Pipelines allow you to specify which existing node properties should be used as features for your machine learning model. These are direct attributes of your nodes.
Let's add 'age' and 'income' properties as features to our pipeline:
CALL gds.pipeline.nodePropertyPrediction.addNodePropertySteps('myChurnPrediction', {
nodeProperties: ['age', 'income']
})
YIELD pipeline
RETURN pipeline.name, pipeline.nodePropertyStepsIntegrating Graph Algorithm Features
The real power of GDS pipelines is integrating graph algorithm results as features. These capture structural insights that simple node properties can't.
We can add PageRank scores as a feature to our pipeline:
CALL gds.pipeline.nodePropertyPrediction.addPageRank('myChurnPrediction', {
maxIterations: 10,
dampingFactor: 0.85,
mutateProperty: 'pageRankScore' // This becomes a feature
})
YIELD pipeline
RETURN pipeline.name, pipeline.pageRankStepsConfiguring the Machine Learning Model
After defining your features, you specify the machine learning model that will perform the prediction. GDS supports various models like Logistic Regression, Random Forest, and GNNs.
Let's add a Logistic Regression model to predict 'isChurned':
CALL gds.pipeline.nodePropertyPrediction.addLogisticRegression('myChurnPrediction', {
targetProperty: 'isChurned', // The node property we want to predict
maxIterations: 100,
penalty: 'l2'
})
YIELD pipeline
RETURN pipeline.name, pipeline.logisticRegressionStepsTraining Your GDS Pipeline
With features and a model defined, you can now train the pipeline. This executes all feature generation steps, then trains the specified ML model on the generated features.
We'll train our pipeline on 'mySocialGraph' and name the resulting model 'churnPredictionModel':
CALL gds.pipeline.nodePropertyPrediction.train('myChurnPrediction', {
modelName: 'churnPredictionModel',
graphName: 'mySocialGraph',
nodeLabels: ['User'],
randomSeed: 42
})
YIELD modelInfo
RETURN modelInfo.modelName AS modelName, modelInfo.trainingMetrics.f1Score.weighted AS f1ScoreMaking Predictions with the Model
Once trained, your model (which is part of the pipeline) can be used to generate predictions on your graph data. You can either stream results or mutate the graph with new properties.
Let's stream predictions for our 'churnPredictionModel':
CALL gds.pipeline.nodePropertyPrediction.predict.stream('churnPredictionModel', {
graphName: 'mySocialGraph',
nodeLabels: ['User'],
topN: 1 // Get the top predicted class
})
YIELD nodeId, predictedProperty, probability
RETURN gds.util.asNode(nodeId).name AS user, predictedProperty, probability
LIMIT 5Managing Your Pipelines
You can list all active pipelines and their configurations using gds.pipeline.list(). When a pipeline is no longer needed, you can remove it with gds.pipeline.drop().
Let's see the pipelines we currently have:
CALL gds.pipeline.list()
YIELD name, type, creationTime
RETURN name, type, creationTimeGDS Pipeline Components
A GDS pipeline combines various steps to create a complete data science workflow. Which of the following are valid components or steps you can add to a GDS Node Property Prediction Pipeline?
Recap: Pipelines for ML
Great job! You've learned how GDS Pipelines provide a structured and reproducible way to integrate graph algorithms with machine learning workflows.
- Pipelines define a sequence of steps.
- They combine feature engineering (from existing properties and graph algorithms) with ML model training.
- They enable efficient prediction and operationalization of graph-based insights.
Keep exploring GDS to unlock more advanced graph analytics!
Questions Fréquemment Posées
La leçon « Pipelines GDS et apprentissage automatique » est-elle gratuite ?
Oui — le texte complet de « Pipelines GDS et apprentissage automatique » est gratuit à lire ici sur le web. Pour la pratiquer de manière interactive (un éditeur de code intégré et un tuteur IA 24/7) et déverrouiller le reste du cours Neo4j Graph Database Fundamentals, passe à CoddyKit PRO. Le cours Neo4j Graph Database Fundamentals comprend 4 leçons au total.
Qu'est-ce que j'apprendrai dans « Pipelines GDS et apprentissage automatique » ?
Découvrez comment créer et gérer des pipelines de science des données appliquée aux graphes dans GDS, en les intégrant à des processus d’apprentissage automatique. Tu pratiques Neo4j Graph Database Fundamentals avec du code pratique que tu exécutes directement dans le navigateur, et un tuteur IA 24/7 répond à tes questions au fur et à mesure que tu avances dans la leçon.
Dois-je avoir de l'expérience pour commencer Neo4j Graph Database Fundamentals ?
Aucune expérience préalable n'est requise. Neo4j Graph Database Fundamentals sur CoddyKit est structuré pour les débutants jusqu'aux apprenants avancés, donc tu peux commencer ici ou depuis le début et avancer à ton rythme. Ceci est la leçon 3 sur 4.
Combien de temps prend la leçon « Pipelines GDS et apprentissage automatique » ?
La plupart des leçons CoddyKit prennent environ 5–10 minutes. Chacune est courte et interactive, tu progresses régulièrement et tu repiques exactement où tu t'es arrêté sur le web et l'app.
Peux-tu écrire et exécuter du code dans cette leçon Neo4j Graph Database Fundamentals ?
Oui. Chaque leçon Neo4j Graph Database Fundamentals inclut un éditeur de code intégré, tu écris et exécutes du vrai code directement dans ton navigateur et tu reçois des retours IA instantanés — aucune configuration locale requise.
Toutes les leçons de ce cours
- Introduction à la bibliothèque GDS
- Exécuter des algorithmes GDS
- Pipelines GDS et apprentissage automatique
- Représentations vectorielles de graphes avec GDS