What it is
Apache Avro is a data serialization system that provides compact, fast, and binary data serialization for Java and other languages. It is widely used in big data ecosystems for storing and transmitting structured data efficiently.
Avro uses JSON-defined schemas to generate Java classes. It supports serialization, deserialization, and schema evolution. Avro can serialize data in a compact binary format or a readable JSON format.
Installation
Add dependency in pom.xml:
<dependency>
<groupId>org.apache.avro</groupId>
<artifactId>avro</artifactId>
<version>1.12.3</version>
</dependency>Getting started
The smallest useful thing you can do with it, and what each part means.
{
"type": "record",
"name": "User",
"namespace": "com.example",
"fields": [
{"name": "name", "type": "string"},
{"name": "age", "type": "int"}
]
}# Using avro-tools
java -jar avro-tools-1.12.3.jar compile schema user.avsc src/main/javaAdvanced usage
Where the library earns its place over a simpler alternative.
User user = User.newBuilder().setName("Alice").setAge(30).build();
ByteArrayOutputStream out = new ByteArrayOutputStream();
DatumWriter<User> writer = new SpecificDatumWriter<>(User.class);
BinaryEncoder encoder = EncoderFactory.get().binaryEncoder(out, null);
writer.write(user, encoder);
encoder.flush();
byte[] serializedData = out.toByteArray();DatumReader<User> reader = new SpecificDatumReader<>(User.class);
BinaryDecoder decoder = DecoderFactory.get().binaryDecoder(serializedData, null);
User deserializedUser = reader.read(null, decoder);
System.out.println(deserializedUser.getName());Schema schema = new Schema.Parser().parse(new File("user.avsc"));
GenericRecord record = new GenericData.Record(schema);
record.put("name", "Bob");
record.put("age", 25);// Add a new optional field in schema without breaking old data
{ "name": "email", "type": ["null", "string"], "default": null }Errors and fixes
The failures you are most likely to hit, and what actually resolves them.
- AvroTypeException
- Occurs when data does not match the schema type. Ensure the data conforms to the defined schema.
- IOException during serialization/deserialization
- Check streams, encoders, decoders, and ensure proper closing of resources.
Best practices
- Use specific Java classes for better type safety when possible.
- Leverage default values in schemas for schema evolution.
- Use Avro binary format for compact storage and JSON format for readability during debugging.
- Integrate with Kafka or Hadoop for streaming and batch data processing.
- Validate schemas before serialization to prevent runtime errors.
Background
Why it exists, and what it was reacting to.
Avro was created as part of the Hadoop project to provide a language-agnostic, schema-based serialization system. It allows schemas to evolve over time, supports rich data structures, and integrates seamlessly with Hadoop, Kafka, and other big data tools. Its compact binary format and dynamic typing make it efficient for data storage and streaming.
