Uploaded image for project: 'Apache Arrow'
  1. Apache Arrow
  2. ARROW-5908

[C#] ArrowStreamWriter doesn't align buffers to 8 bytes

Attach filesAttach ScreenshotVotersWatch issueWatchersCreate sub-taskLinkCloneUpdate Comment AuthorReplace String in CommentUpdate Comment VisibilityDelete Comments
    XMLWordPrintableJSON

Details

    Description

      When writing RecordBatches using ArrowStreamWriter, if the ArrowBuffers being written aren't all 8 byte aligned, the serialized RecordBatch won't conform to the Arrow specification. This leads to other languages' readers to throw an error when reading Arrow streams written by the C# writer.

      For example, if reading the stream from Python or C++, an error is raised here: 

      https://github.com/apache/arrow/blob/f77c3427ca801597b572fb197b92b0133269049b/cpp/src/arrow/ipc/reader.cc#L107-L110

      A similar error is raised when Java tries to read the stream.

      We should be ensuring that the buffers being written to the stream are padded to 8 bytes, no matter their length, as specified in https://arrow.apache.org/docs/format/Layout.html#requirements-goals-and-non-goals

       

      • It is required to have all the contiguous memory buffers in an IPC payload aligned at 8-byte boundaries. In other words, each buffer must start at an aligned 8-byte offset. Additionally, each buffer should be padded to a multiple of 8 bytes.

      Attachments

        Activity

          This comment will be Viewable by All Users Viewable by All Users
          Cancel

          People

            eerhardt Eric Erhardt
            eerhardt Eric Erhardt
            Votes:
            0 Vote for this issue
            Watchers:
            2 Start watching this issue

            Dates

              Created:
              Updated:
              Resolved:

              Time Tracking

                Estimated:
                Original Estimate - Not Specified
                Not Specified
                Remaining:
                Remaining Estimate - 0h
                0h
                Logged:
                Time Spent - 1h 40m
                1h 40m

                Slack

                  Issue deployment