Parquet file is an hdfs file that must include the metadata for the file. This allows splitting columns into multiple files, as well as having a single metadata file reference multiple parquet files. The metadata includes the schema for the data stored in the file.
How do I create a schema for a parquet file?
To generate the schema of the parquet sample data, do the following:
- Log in to the Haddop/Hive box.
- It generates the schema in the stdout as follows: -------------- [ ~]# parquet-tools schema abc.parquet. message hive_schema { ...
- Copy this schema to a file with . parquet/. par extension.
Does parquet support schema evolution?
Schema Merging
Like Protocol Buffer, Avro, and Thrift, Parquet also supports schema evolution. Users can start with a simple schema, and gradually add more columns to the schema as needed. In this way, users may end up with multiple Parquet files with different but mutually compatible schemas.