Uploaded image for project: 'Slider'
  1. Slider
  2. SLIDER-823 Über-JIRA : placement phase 3
  3. SLIDER-870

use timeline server as a historical source of failure information

    XMLWordPrintableJSON

Details

    • Sub-task
    • Status: Open
    • Major
    • Resolution: Unresolved
    • Slider 0.80
    • Slider 1.0.0
    • appmaster, client
    • None

    Description

      We lose failure history when an AM dies; this hurts reporting and doesn't allow the collection of long-term statistics.

      We can use the timeline server for this information, saving events on failure, then querying it on AM restart to rebuild that history & re-use it in decision making.

      They can also be presented to the user in (a) the web UI and (b) from the command line —even while a cluster is not running.

      Finally, stats on node failures could be aggregated across applications, possibly even across users. This would identify hotspots for node unreliability.

      Attachments

        Activity

          People

            Unassigned Unassigned
            stevel@apache.org Steve Loughran
            Votes:
            0 Vote for this issue
            Watchers:
            3 Start watching this issue

            Dates

              Created:
              Updated: