Uploaded image for project: 'Apache Arrow'
  1. Apache Arrow
  2. ARROW-17008

[R] Parquet Snappy Compression Fails for Integer Type Data

    XMLWordPrintableJSON

Details

    • Bug
    • Status: Open
    • Major
    • Resolution: Unresolved
    • 8.0.0
    • None
    • Documentation, R
    • None
    • R4.2.1 Ubuntu 22.04 x86_64
      R4.1.2 Ubuntu 22.04 Aarch64

    Description

      Snappy compression is not working when writing to parquet for integer type data.

      E.g. compare file sizes for:

      write_parquet(data.frame(x = 1:1e6), "snappy.parquet", compression = "snappy")
      write_parquet(data.frame(x = 1:1e6), "uncomp.parquet", compression = "uncompressed")
      

      whereas for double:

      write_parquet(data.frame(x = as.double(1:1e6)), "snappyd.parquet", compression = "snappy")
      write_parquet(data.frame(x = as.double(1:1e6)), "uncompd.parquet", compression = "uncompressed")
      

      I have inspected the integer files using parquet-tools and compression level shows as 0%. Needless to say, I can achieve compression using Spark (sparklyr) etc.

      Thanks.

      Attachments

        Activity

          People

            Unassigned Unassigned
            charliegao Charlie Gao
            Votes:
            0 Vote for this issue
            Watchers:
            3 Start watching this issue

            Dates

              Created:
              Updated: