[IMPALA-2753] Investigate performance gains for adding random prefix to file name - ASF JIRA

XML

Word

Printable

JSON

Details

Type: Sub-task
Status: Resolved
Priority: Minor
Resolution: Won't Fix
Affects Version/s: Impala 2.5.0
Fix Version/s: None
Component/s: Perf Investigation
Labels:
- s3

Target Version:

Product Backlog

Description

I noticed which is not directly related to Impala is that the file naming convention HDFS produces is the anti pattern of what S3 recommends.
If we do a trick with the naming convention we can one up Hive when running on S3.

examplebucket/2013-26-05-15-00-00/cust1234234/photo1.jpg
examplebucket/2013-26-05-15-00-00/cust3857422/photo2.jpg
examplebucket/2013-26-05-15-00-00/cust8474937/photo2.jpg
examplebucket/2013-26-05-15-00-00/cust1248473/photo3.jpg
...
examplebucket/2013-26-05-15-00-01/cust1248473/photo4.jpg
examplebucket/2013-26-05-15-00-01/cust1248473/photo5.jpg
examplebucket/2013-26-05-15-00-01/cust1248473/photo6.jpg
examplebucket/2013-26-05-15-00-01/cust1248473/photo7.jpg    
...

The sequence pattern in the key names introduces a performance problem. To understand the issue, let’s look at how Amazon S3 stores key names.

Amazon S3 maintains an index of object key names in each AWS region. Object keys are stored lexicographically across multiple partitions in the index. That is, Amazon S3 stores key names in alphabetical order. The key name dictates which partition the key is stored in. Using a sequential prefix, such as timestamp or an alphabetical sequence, increases the likelihood that Amazon S3 will target a specific partition for a large number of your keys, overwhelming the I/O capacity of the partition. If you introduce some randomness in your key name prefixes, the key names, and therefore the I/O load, will be distributed across more than one partition.

If you anticipate that your workload will consistently exceed 100 requests per second, you should avoid sequential key names. If you must use sequential numbers or date and time patterns in key names, add a random prefix to the key name. The randomness of the prefix more evenly distributes key names across multiple index partitions. Examples of introducing randomness are provided later in this topic.

http://docs.aws.amazon.com/AmazonS3/latest/dev/request-rate-perf-considerations.html

Attachments

Issue Links

relates to

IMPALA-10403 Implement ObjectStoreLocationProvider for Iceberg tables

Open

Activity

People

Assignee:: Unassigned

Reporter:: Mostafa Mokhtar

Votes:: 0 Vote for this issue

Watchers:: 4 Start watching this issue

Dates

Created:: 08/Dec/15 20:37

Updated:: 22/Dec/20 17:33

Resolved:: 22/Dec/20 17:33