[NUTCH-356] Plugin repository cache can lead to memory leak - ASF JIRA

XML

Word

Printable

JSON

Details

Type: Bug
Status: Closed
Priority: Major
Resolution: Fixed
Affects Version/s: 0.8
Fix Version/s: 2.3, 1.8
Component/s: None
Labels:
None

Patch Info:

Patch Available

Description

While I was trying to solve a problem I reported a while ago (see Nutch-314), I found out that actually the problem was related to the plugin cache used in class PluginRepository.java.
As I said in Nutch-314, I think I somehow 'force' the way nutch is meant to work, since I need to frequently submit new urls and append their contents to the index; I don't (and I can't) have an urls.txt file with all urls I'm going to fetch, but I recreate it each time a new url is submitted.
Thus, I think in the majority of times you won't have problems using nutch as-is, since the problem I found occours only if nutch is used in a way similar to the one I use.
To simplify your test I'm attaching a class that performs something similar to what I need. It fetches and index some sample urls; to avoid webmasters complaints I left the sample urls list empty, so you should modify the source code and add some urls.
Then you only have to run it and watch your memory consumption with top. In my experience I get an OutOfMemoryException after a couple of minutes, but it clearly depends on your heap settings and on the plugins you are using (I'm using 'protocol-file|protocol-http|parse-(rss|html|msword|pdf|text)|language-identifier|index-(basic|more)|query-(basic|more|site|url)|urlfilter-regex|summary-basic|scoring-opic').

The problem is bound to the PluginRepository 'singleton' instance, since it never get released. It seems that some class maintains a reference to it and this class is never released since it is cached somewhere in the configuration.

So I modified the PluginRepository's 'get' method so that it never uses the cache and always returns a new instance (you can find the patch in attachment). This way the memory consumption is always stable and I get no OOM anymore.
Clearly this is not the solution, since I guess there are many performance issues involved, but for the moment it works.

Attachments

- Sort By Name
- Sort By Date
- Ascending
- Descending

ASF.LICENSE.NOT.GRANTED--NutchTest.java
21/Aug/06 11:59
4 kB
Enrico Triolo
ASF.LICENSE.NOT.GRANTED--patch.txt
21/Aug/06 11:59
0.9 kB
Enrico Triolo
cache_classes.patch
24/Jun/07 19:05
6 kB
Dogacan Guney
NUTCH-356-trunk.patch
02/Jan/14 11:36
6 kB
Markus Jelsma

Issue Links

is part of

NUTCH-844 Improve NutchConfiguration

Closed

Activity

People

Assignee:: Markus Jelsma

Reporter:: Enrico Triolo

Votes:: 3 Vote for this issue

Watchers:: 7 Start watching this issue

Dates

Created:: 21/Aug/06 11:59

Updated:: 01/May/14 06:23

Resolved:: 24/Jan/14 13:22