Schema Evolution Strategies
Understand techniques for evolving Protobuf schemas without breaking existing clients or services.
Schema Evolution Strategies is a free gRPC & High Performance APIs lesson on CoddyKit — lesson 2 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the gRPC & High Performance APIs learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.
Why Schema Evolution Matters
In distributed systems, services and clients often need to communicate using a defined data format, like Protocol Buffers (Protobuf).
Over time, these data structures need to change. Maybe you need to add a new field, remove an old one, or change a type.
Schema evolution is the art of changing your data definitions without breaking existing, older versions of your services or clients. It's crucial for maintaining compatibility in dynamic environments.
The Challenge of Compatibility
When you update a schema, you face two main challenges:
- Backward Compatibility: Can an older client still communicate with a newer server? The server must understand the old client's requests.
- Forward Compatibility: Can a newer client still communicate with an older server? The server must gracefully ignore new fields it doesn't understand.
Breaking compatibility can lead to service outages and difficult deployments.
Protobuf's Key: Field Numbers
Unlike JSON, where field names are used for identification, Protobuf uses unique field numbers to identify fields in your messages.
These numbers are critical for compatibility. When a message is serialized, only the field numbers and their values are stored, not the field names.
This means:
- Field numbers must be unique within a message.
- Once assigned, a field number should never change.
- Once assigned, a field number should never be reused, even if the field is removed.
Strategy 1: Adding New Fields
Adding new fields is generally safe, provided you follow these rules:
- Assign a new, unused field number.
- Make the new field
optional(orrepeated,mapinproto3).
Old clients will simply ignore the new field. New clients communicating with old servers will use the field's default value if it's not present.
Try running this Java code example to see how a Protobuf-generated message handles a new field:
import com.google.protobuf.InvalidProtocolBufferException;
import com.google.protobuf.util.JsonFormat;
// Assume these classes are generated from .proto files:
// Original: message User { string name = 1; }
// Evolved: message User { string name = 1; int32 age = 2; }
// We'll simulate the User class for demonstration purposes.
class User {
private final String name;
private final int age;
private User(Builder builder) {
this.name = builder.name;
this.age = builder.age;
}
public String getName() { return name; }
public int getAge() { return age; }
public static Builder newBuilder() { return new Builder(); }
public static class Builder {
private String name = "";
private int age = 0; // Default value for new field
public Builder setName(String name) { this.name = name; return this; }
public Builder setAge(int age) { this.age = age; return this; }
public User build() { return new User(this); }
}
@Override
public String toString() { return "User{name='" + name + "', age=" + age + "}"; }
}
public class AddFieldEvolution {
public static void main(String[] args) {
// Simulate an old client sending data (unaware of 'age')
User oldClientUser = User.newBuilder()
.setName("Alice")
.build();
System.out.println("Old client sends: " + oldClientUser);
// Simulate a new server receiving this data.
// The 'age' field will correctly default to 0.
System.out.println("New server receives (age): " + oldClientUser.getAge());
// Simulate a new client sending data (aware of 'age')
User newClientUser = User.newBuilder()
.setName("Bob")
.setAge(30)
.build();
System.out.println("New client sends: " + newClientUser);
// Simulate an old server receiving this data.
// It will simply ignore the 'age' field.
System.out.println("Old server receives (name only): " + newClientUser.getName());
}
}Strategy 2: Removing Fields
You should never truly delete a field number, as this could lead to data corruption if the number is reused later.
Instead, mark fields as deprecated and reserved:
- Use the
deprecated = trueoption to signal that the field should no longer be used. Compilers will issue warnings. - Use the
reservedkeyword to prevent future assignment of specific field numbers or names. This ensures the number is never accidentally reused.
Here's how you'd mark a field as deprecated and reserve its number:
syntax = "proto3";
package evolution;
message OldMessage {
string id = 1;
// This field is deprecated and should not be used.
string old_data = 2 [deprecated = true];
string new_data = 3;
// Reserve field number 2 and the name 'old_data'
// to prevent accidental reuse in the future.
reserved 2;
reserved "old_data";
}Strategy 3: Renaming Fields
Remember, Protobuf identifies fields by their field numbers, not their names. So, simply changing a field's name in the .proto file is a compatible change.
However, if you also need to change the field number, this is effectively a 'remove' followed by an 'add' operation. In such cases:
- Mark the old field number as
reserved. - Add a new field with the new name and a new, unused field number.
This ensures that old clients/servers don't get confused by conflicting field numbers.
Strategy 4: Changing Field Types
Changing a field's type is often not backward or forward compatible and should be done with extreme caution.
Some safe changes:
int32toint64(values will be truncated if read by old client).uint32touint64.
Unsafe changes (will break compatibility):
int32tostring.int32tofixed32.- Any change involving
enum,message, orbytesto other types.
If an unsafe type change is unavoidable, treat it as removing the old field and adding a new one with a new number.
Strategy 5: Evolving Enums
Enums are represented as integers. Adding new values to an enum is generally safe, but follow these rules:
- Always add new enum values to the end of the list.
- Assign a new, unused integer value.
- Never change the numeric value of an existing enum member.
Old clients encountering a new enum value will typically see its integer representation, which they might not handle gracefully if they expect only known values. Always include a 0 value as the first enum member for compatibility.
syntax = "proto3";
package evolution;
message StatusUpdate {
Status current_status = 1;
}
enum Status {
UNKNOWN = 0;
PENDING = 1;
PROCESSING = 2;
// New status added (safe)
COMPLETED = 3;
// Another new status (safe)
FAILED = 4;
}Strategy 6: Evolving Oneof Fields
A oneof field means that at most one of the fields within the oneof group can be set at a time.
Evolving oneof fields follows similar rules:
- Adding new fields to a
oneofis compatible. Assign a new, unused field number. Old clients will ignore these new cases. - Removing fields from a
oneofrequires deprecating and reserving the field number, just like regular fields.
Be careful when changing existing fields within a oneof, as this can affect compatibility.
Quick Check: Schema Rules
Which of the following actions is generally considered unsafe and likely to break Protobuf compatibility?
Recap: Safe Schema Evolution
Congratulations! You've learned the key strategies for evolving your Protobuf schemas safely:
- Field Numbers: Are paramount and must be unique and stable. Never change or reuse them.
- Adding Fields: Always assign new numbers; new fields are ignored by old clients.
- Removing Fields: Deprecate and reserve field numbers to prevent future reuse.
- Renaming Fields: Only change the name, not the number, or treat as remove/add.
- Type Changes: Mostly unsafe; avoid or treat as remove/add.
- Enums: Add new values to the end, never change existing numbers.
By following these guidelines, you can ensure your gRPC services remain compatible as they evolve.
Frequently asked questions
Is the “Schema Evolution Strategies” lesson free?
Yes — the full text of “Schema Evolution Strategies” is free to read here on the web, and the gRPC & High Performance APIs course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the gRPC & High Performance APIs course, upgrade to CoddyKit PRO.
What will I learn in “Schema Evolution Strategies”?
Understand techniques for evolving Protobuf schemas without breaking existing clients or services. You practise gRPC & High Performance APIs with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.
Do I need any experience to start gRPC & High Performance APIs?
No prior experience is required. gRPC & High Performance APIs on CoddyKit is structured for beginners through advanced learners; this is — lesson 2 of 4, so you can start here or from the beginning and move at your own pace.
How long does the “Schema Evolution Strategies” lesson take?
Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.
Can I write and run code in this gRPC & High Performance APIs lesson?
Yes. Every gRPC & High Performance APIs lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.
All lessons in this course
- Protobuf Best Practices
- Schema Evolution Strategies
- Custom Protobuf Options
- Oneof, Maps & Well-Known Types