Commit Graph
11 Commits
Author SHA1 Message Date
Максименко Никита ВладимировичandClaude Sonnet 5 f553e79c75 refactor: dictionary lookups from O(dictionary size) to O(word length)
isWellKnown/isKnownCity scanned the whole stem list (60, then 1111+
after the city dictionary) with startsWith on every call. Reversed the
check: try decreasing-length prefixes of the input word against a
HashSet of stems instead — same result, but bounded by word length,
not dictionary size.

Also pulls the vowel-stripping declension trick (shared verbatim
between NameDictionary and ToponymDictionary since the city dictionary
was added) into one Declension helper, and extracts the {{TYPE:value}}
benchmark-line parser — duplicated between BenchmarkTest and
LargeTextTest since the large-text test was added — into
BenchmarkFixtures.

Verified: same 125 tests pass, benchmark F1 numbers unchanged, native
image under 0.5 CPU/250MB shows no regression at 1000/1800/3000 RPS.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-21 21:51:30 +03:00
Максименко Никита ВладимировичandClaude Sonnet 5 e00062ea9f test: real Alfa-Bank addresses from the CBR registry; fix two gaps they found
alfabank.ru itself is unreachable from this environment, so addresses
came from the official regulator registry instead (cbr.ru/finorg/foinfo,
"Alfa-Bank subdivisions") — 981 addresses across 509 cities, one per
region sampled into benchmark-bank-context.txt (60 lines) as the
office-address trap the spec calls out explicitly.

Running them surfaced two real precision bugs, not just more test
data:

- ORGANISATION_NEARBY's veto window (80 chars) was too narrow for
  official addresses that include a region name before the city —
  "Отделение ... Республика Бурятия, г. Северобайкальск, ..." puts
  the house number 80+ chars from the anchor. Widened to 150.

- The two dictionary-only FIO rules (surname + capitalized word,
  no role word) had no address-context veto at all: "Великие Луки"
  matched as the given name "Лука" plus a stray word, "Богдана
  Хмельницкого" as a person because streets named after people are
  syntactically identical to actual names. Added the same
  ORGANISATION_NEARBY veto the address rules already use.

False-positive rate on the 60-address sample: 40.5% before, 0% after.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-21 21:29:57 +03:00
Максименко Никита ВладимировичandClaude Sonnet 5 495cf08a05 feat: detect account number, BIK, card expiry, OGRN(IP), KPP, income, biometrics
Types the spec asks for beyond the payment card itself: settlement
account (20 digits), BIK, card expiry date (kept separate from
birth date), OGRN/OGRNIP (negative lookahead so OGRNIP isn't half-
swallowed by the OGRN rule), KPP, income/salary amounts, and mentions
of biometric enrollment (ЕБС, voice/fingerprint templates).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-21 21:29:57 +03:00
Максименко Никита ВладимировичandClaude Sonnet 5 e52e9550e6 feat: extend requireCompanion to BIRTH_PLACE and ADDRESS_COUNTRY
Same reasoning as the existing CVV/PIN/DATE entries: a bare mention
("Экскурсия в Нижний Новгород", "цены выросли в Казахстане") isn't
personal data on its own, only alongside some other PD that ties it to
a specific person. The mechanism was already generic — this only adds
two more types to the default set.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-21 21:29:57 +03:00
Максименко Никита ВладимировичandClaude Sonnet 5 7dd66d45bd feat: expandable denylist + hot-reload, fix declension for -a surnames
Adds current public figures (Nabiullina, Gref, Putin, ...) so news-
style mentions near CB/banking topics aren't treated as client PII.

More importantly, isWellKnown() only worked for consonant-ending
surnames: "Пушкина" starts with "Пушкин" so it matched, but "Толстого"
doesn't start with "Толстой" (adjectival declension replaces the
ending, doesn't extend it), and neither does any oblique case of a
feminine -a surname ("Набиуллиной" vs "Набиуллина"). Strips the
trailing vowel at load time, same trick already used for given names.

Also lets the list grow without a rebuild: an optional external file
(pdguard.well-known-file, default config/well-known.txt) is merged on
top of the bundled list and re-read on change, mirroring how
SystemsConfig watches systems.json.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-21 21:29:57 +03:00
Максименко Никита ВладимировичandClaude Sonnet 5 f7eee7577c feat: validate ADDRESS_CITY against a real city dictionary
The rule matched any capitalized word after "г."/"город" — no check
that it's an actual place. ToponymDictionary checks the match against
1111 Russian cities (pensnarik/russian-cities) plus CIS capitals,
matching by stem so declined forms work ("Москве" against "Москва").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-21 21:29:57 +03:00
Максименко Никита ВладимировичandClaude Sonnet 5 a758c5ef05 test: independent generated benchmark + tests on large mixed text
benchmark-generated.txt: a fresh dataset (not used to tune the rules)
covering every PD type from the spec plus format variations — date
order, "серия ... номер ...", register-independence. Surfaced two real
gaps worth tracking: FIO detection needs a capitalized first letter,
so "клиент иванова мария петровна" (all lowercase) isn't found; the
NER cascade flags plain capitalized nouns ("Портрет", "Рим") as names
when the base rules alone don't.

LargeTextTest: the existing large-text test just repeated one
email+card sentence 4000 times, so 27 of 28 PD types never ran at
scale. Replaces it with ~400,000 characters built from shuffled lines
of the new dataset — recall on it holds at 0.970, matching small-scale
numbers, and the NER cascade stays sub-second on the whole thing.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-21 21:29:57 +03:00
Максименко Никита ВладимировичandClaude Sonnet 5 832738891c feat: self-tuning concurrency limit instead of a fixed Semaphore
Under 0.5-CPU containers, a static Semaphore(2000) never tripped —
latency ballooned to 1.5-2s instead of the service answering 429.
Runtime.availableProcessors() can't help pick a number either: it
ignores the cgroups --cpus quota and reports full host cores.

AdaptiveConcurrencyLimiter reacts to observed latency instead of
guessing capacity: starts at min-concurrent, grows by one per
adjustment window when latency stays under target, halves it the
moment it doesn't. Adjustment is gated by wall-clock time, not by
request count — an earlier per-request version let the limit race to
the ceiling in milliseconds under high RPS, before any real overload
had a chance to show up in the samples.

Verified under load (native image, 250MB/0.5 CPU): p50 latency at 3x
overload dropped from ~1.3s to under 4ms; normal-load p95 unaffected.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-21 21:29:57 +03:00
dakocha3 c7e5b02362 Fix for jvm Dockerfile.jvm 2026-09-21 18:59:00 +03:00
dakocha3 13f5a93533 Add model 2026-09-21 18:11:34 +03:00
dakocha3 309188d191 Init 2026-09-21 17:40:25 +03:00