Uploaded image for project: 'Spark'
  1. Spark
  2. SPARK-17210

sparkr.zip is not distributed to executors when run sparkr in RStudio

    XMLWordPrintableJSON

Details

    • Bug
    • Status: Resolved
    • Major
    • Resolution: Fixed
    • 2.0.0
    • 2.0.1, 2.1.0
    • SparkR
    • None

    Description

      Here's the code to reproduce this issue.

      Sys.setenv(SPARK_HOME="/Users/jzhang/github/spark")
      .libPaths(c(file.path(Sys.getenv(), "R", "lib"), .libPaths()))
      library(SparkR)
      sparkR.session(master="yarn-client", sparkConfig = list(spark.executor.instances="1"))
      df <- as.DataFrame(mtcars)
      head(df)
      

      And this is the exception in executor log.

      16/08/24 15:33:45 INFO BufferedStreamThread: Fatal error: cannot open file '/Users/jzhang/Temp/hadoop_tmp/nm-local-dir/usercache/jzhang/appcache/application_1471846125517_0022/container_1471846125517_0022_01_000002/sparkr/SparkR/worker/daemon.R': No such file or directory
      16/08/24 15:33:55 ERROR Executor: Exception in task 0.0 in stage 3.0 (TID 6)
      java.net.SocketTimeoutException: Accept timed out
          at java.net.PlainSocketImpl.socketAccept(Native Method)
          at java.net.AbstractPlainSocketImpl.accept(AbstractPlainSocketImpl.java:404)
          at java.net.ServerSocket.implAccept(ServerSocket.java:545)
          at java.net.ServerSocket.accept(ServerSocket.java:513)
          at org.apache.spark.api.r.RRunner$.createRWorker(RRunner.scala:367)
          at org.apache.spark.api.r.RRunner.compute(RRunner.scala:69)
          at org.apache.spark.api.r.BaseRRDD.compute(RRDD.scala:49)
          at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)
          at org.apache.spark.rdd.RDD.iterator(RDD.scala:283)
          at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)
          at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)
          at org.apache.spark.rdd.RDD.iterator(RDD.scala:283)
          at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)
          at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)
          at org.apache.spark.rdd.RDD.iterator(RDD.scala:283)
          at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)
          at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)
          at org.apache.spark.rdd.RDD.iterator(RDD.scala:283)
          at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)
          at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)
          at org.apache.spark.rdd.RDD.iterator(RDD.scala:283)
          at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:70)
          at org.apache.spark.scheduler.Task.run(Task.scala:86)
          at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)
          at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)
          at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)
          at java.lang.Thread.run(Thread.java:745)
      

      Attachments

        Issue Links

          Activity

            People

              zjffdu Jeff Zhang
              zjffdu Jeff Zhang
              Votes:
              1 Vote for this issue
              Watchers:
              5 Start watching this issue

              Dates

                Created:
                Updated:
                Resolved: