Man are you fast!
not really, i've been working it for a while but since someone asked i figure i would create the issue.
testing isnt done, but english, french, portuguese I think are ok.
the others need a lot of tests and probably have bugs.
Does the English one deal with women/ woman and foci / focus type stuff?
Nope, the english one is the Harman "s-stemming" algorithm.
its very simple:
if final is '-ies' but not '-eies' or '-aies' then
replace '-ies' by '-y', return;
if final is '-es' but not '-aes', '-ees' or '-oes' then
replace '-es' by '-e', return;
if final is '-s' but not '-us' or '-ss' then
For special cases like you mentioned (if you want them), i would recommend adding these customizations yourself
as documented here: http://wiki.apache.org/solr/LanguageAnalysis#Customizing_Stemming
just make a tab-separated file of words-stems and put a StemmerOverrideFilter(Factory) before the stemmer in the stream.
I think this alone provides a lot of flexibility. if it isn't enough, then i think these stemmers are much simpler to modify if you wanted to go that route also