Details
Description
Nutch would benefit from having a DB with per-host/domain/TLD information. For instance, Nutch could detect hosts that are timing out, store information about that in this DB. Segment/fetchlist Generator could then skip such hosts, so they don't slow down the fetch job. Another good use for such a DB is keeping track of various host scores, e.g. spam score.
From the recent thread on nutch-user@lucene:
Otis asked:
> While we are at it, how would one go about implementing this DB, as far as its structures go?
Andrzej said:
The easiest I can imagine is to use something like <Text, MapWritable>.
This way you could store arbitrary information under arbitrary keys.
I.e. a single database then could keep track of aggregate statistics at
different levels, e.g. TLD, domain, host, ip range, etc. The basic set
of statistics could consist of a few predefined gauges, totals and averages.