[SPARK-32953] Lower memory usage in toPandas with Arrow self_destruct - ASF JIRA

XML

Word

Printable

JSON

Details

Type: Improvement
Status: Resolved
Priority: Major
Resolution: Fixed
Affects Version/s: 3.0.1
Fix Version/s: 3.2.0
Component/s: PySpark
Labels:
None

Description

As described on the mailing list:
http://apache-spark-developers-list.1001551.n3.nabble.com/DISCUSS-Reducing-memory-usage-of-toPandas-with-Arrow-quot-self-destruct-quot-option-td30149.html https://lists.apache.org/thread.html/r581d7c82ada1c2ac3f0584615785cc60cf5ac231e1f29737d3a6569f%40%3Cdev.spark.apache.org%3E

toPandas() can as much as double memory usage as both Arrow and Pandas retain a copy of a dataframe in memory during the conversion. Arrow >= 0.16 offers a self_destruct mode that avoids this with some caveats.

Attachments

Issue Links

is related to

SPARK-34463 toPandas failed with error: buffer source array is read-only when Arrow with self-destruct is enabled

Resolved

links to

[Github] Pull Request #29818 (lidavidm)

Activity

People

Assignee:: David Li

Reporter:: David Li

Votes:: 0 Vote for this issue

Watchers:: 3 Start watching this issue

Dates

Created:: 21/Sep/20 13:18

Updated:: 09/Aug/21 01:59

Resolved:: 10/Feb/21 17:59