isWellKnown/isKnownCity scanned the whole stem list (60, then 1111+
after the city dictionary) with startsWith on every call. Reversed the
check: try decreasing-length prefixes of the input word against a
HashSet of stems instead — same result, but bounded by word length,
not dictionary size.
Also pulls the vowel-stripping declension trick (shared verbatim
between NameDictionary and ToponymDictionary since the city dictionary
was added) into one Declension helper, and extracts the {{TYPE:value}}
benchmark-line parser — duplicated between BenchmarkTest and
LargeTextTest since the large-text test was added — into
BenchmarkFixtures.
Verified: same 125 tests pass, benchmark F1 numbers unchanged, native
image under 0.5 CPU/250MB shows no regression at 1000/1800/3000 RPS.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds current public figures (Nabiullina, Gref, Putin, ...) so news-
style mentions near CB/banking topics aren't treated as client PII.
More importantly, isWellKnown() only worked for consonant-ending
surnames: "Пушкина" starts with "Пушкин" so it matched, but "Толстого"
doesn't start with "Толстой" (adjectival declension replaces the
ending, doesn't extend it), and neither does any oblique case of a
feminine -a surname ("Набиуллиной" vs "Набиуллина"). Strips the
trailing vowel at load time, same trick already used for given names.
Also lets the list grow without a rebuild: an optional external file
(pdguard.well-known-file, default config/well-known.txt) is merged on
top of the bundled list and re-read on change, mirroring how
SystemsConfig watches systems.json.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>