[SPARK-19728] PythonUDF with multiple parents shouldn't be pushed down when used as a predicate - ASF JIRA

XML

Word

Printable

JSON

Details

Type: Bug
Status: Closed
Priority: Major
Resolution: Fixed
Affects Version/s: 2.0.0, 2.1.0
Fix Version/s: 2.2.0
Component/s: PySpark, SQL
Labels:
None

Description

Prior to Spark 2.0 it was possible to use Python UDF output as a predicate:

from pyspark.sql.functions import udf
from pyspark.sql.types import BooleanType

df1 = sc.parallelize([(1, ), (2, )]).toDF(["col_a"])
df2 = sc.parallelize([(2, ), (3, )]).toDF(["col_b"])
pred = udf(lambda x, y: x == y, BooleanType())

df1.join(df2).where(pred("col_a", "col_b")).show()

In Spark 2.0 this is no longer possible:

spark.conf.set("spark.sql.crossJoin.enabled", True)
df1.join(df2).where(pred("col_a", "col_b")).show()

## ...
## Py4JJavaError: An error occurred while calling o731.showString.
: java.lang.RuntimeException: Invalid PythonUDF <lambda>(col_a#132L, col_b#135L), requires attributes from more than one child.
## ...

Attachments

Issue Links

duplicates

SPARK-18589 persist() resolves "java.lang.RuntimeException: Invalid PythonUDF <lambda>(...), requires attributes from more than one child"

Resolved

Activity

People

Assignee:: Unassigned

Reporter:: Maciej Szymkiewicz

Votes:: 2 Vote for this issue

Watchers:: 3 Start watching this issue

Dates

Created:: 24/Feb/17 14:32

Updated:: 15/May/19 14:59

Resolved:: 23/Mar/17 14:04