[DRILL-4530] Improve metadata cache performance for queries with single partition - ASF JIRA

XML

Word

Printable

JSON

Details

Type: Improvement
Status: Closed
Priority: Major
Resolution: Fixed
Affects Version/s: 1.6.0
Fix Version/s: 1.8.0
Component/s: Query Planning & Optimization
Labels:
None

Description

Consider two types of queries which are run with Parquet metadata caching:

query 1:
SELECT col FROM  `A/B/C`;

query 2:
SELECT col FROM `A` WHERE dir0 = 'B' AND dir1 = 'C';

For a certain dataset, the query1 elapsed time is 1 sec whereas query2 elapsed time is 9 sec even though both are accessing the same amount of data. The user expectation is that they should perform roughly the same. The main difference comes from reading the bigger metadata cache file at the root level 'A' for query2 and then applying the partitioning filter. query1 reads a much smaller metadata cache file at the subdirectory level.

Attachments

Issue Links

Dependent

ZOOKEEPER-704 GSoC 2010: Read-Only Mode

Open

links to

GitHub Pull Request #519

Activity

People

Assignee:: Aman Sinha

Reporter:: Aman Sinha

Reviewer:: Rahul Kumar Challapalli

Votes:: 0 Vote for this issue

Watchers:: 4 Start watching this issue

Dates

Created:: 23/Mar/16 00:51

Updated:: 14/Apr/21 05:57

Resolved:: 19/Jul/16 15:20