[HIVE-21492] VectorizedParquetRecordReader can't to read parquet file generated using thrift/custom tool - ASF JIRA

XML

Word

Printable

JSON

Details

Type: Bug
Status: Closed
Priority: Major
Resolution: Fixed
Affects Version/s: None
Fix Version/s: 4.0.0-alpha-1
Component/s: Parquet
Labels:
None

Description

Taking an example of a parquet table having array of integers as below.

CREATE EXTERNAL TABLE ( list_of_ints` array<int>)
STORED AS PARQUET 
LOCATION '{location}';

Parquet file generated using hive will have schema for Type as below:

group list_of_ints (LIST) { repeated group bag { optional int32 array;\n};\n}

Parquet file generated using thrift or any custom tool (using org.apache.parquet.io.api.RecordConsumer)

may have schema for Type as below:

required group list_of_ints (LIST) { repeated int32 list_of_tuple}

VectorizedParquetRecordReader handles only parquet file generated using hive. It throws the following exception when parquet file generated using thrift is read because of the changes done as part of ~~HIVE-18553~~ .

Caused by: java.lang.ClassCastException: repeated int32 list_of_ints_tuple is not a group
 at org.apache.parquet.schema.Type.asGroupType(Type.java:207)
 at org.apache.hadoop.hive.ql.io.parquet.vector.VectorizedParquetRecordReader.getElementType(VectorizedParquetRecordReader.java:479)
 at org.apache.hadoop.hive.ql.io.parquet.vector.VectorizedParquetRecordReader.buildVectorizedParquetReader(VectorizedParquetRecordReader.java:532)
 at org.apache.hadoop.hive.ql.io.parquet.vector.VectorizedParquetRecordReader.checkEndOfRowGroup(VectorizedParquetRecordReader.java:440)
 at org.apache.hadoop.hive.ql.io.parquet.vector.VectorizedParquetRecordReader.nextBatch(VectorizedParquetRecordReader.java:401)
 at org.apache.hadoop.hive.ql.io.parquet.vector.VectorizedParquetRecordReader.next(VectorizedParquetRecordReader.java:353)
 at org.apache.hadoop.hive.ql.io.parquet.vector.VectorizedParquetRecordReader.next(VectorizedParquetRecordReader.java:92)
 at org.apache.hadoop.hive.ql.io.HiveContextAwareRecordReader.doNext(HiveContextAwareRecordReader.java:365)

I have done a small change to handle the case where the child type of group type can be PrimitiveType.

Attachments

- Sort By Name
- Sort By Date
- Ascending
- Descending

HIVE-21492.patch
22/Mar/19 16:02
1 kB
Ganesha Shreedhara
HIVE-21492.2.patch
01/Apr/20 02:57
4 kB
Ganesha Shreedhara
HIVE-21492.3.patch
01/Apr/20 07:28
4 kB
Ganesha Shreedhara
HIVE-21492.4.patch
03/Apr/20 03:31
4 kB
Ganesha Shreedhara
HIVE-21492.5.patch
05/Apr/20 02:47
4 kB
Ganesha Shreedhara

Activity

People

Assignee:: Ganesha Shreedhara

Reporter:: Ganesha Shreedhara

Votes:: 0 Vote for this issue

Watchers:: 3 Start watching this issue

Dates

Created:: 22/Mar/19 15:59

Updated:: 17/Nov/22 08:51

Resolved:: 07/Apr/20 15:26