Back to Blog

Scraping REGON GUS, RDF KRS, SUDOP, EKW and Polish Registries with Mobile Proxies

By: Proxy PolandPublished: Last updated:

Real client traffic hits wyszukiwarkaregon.stat.gov.pl first, then RDF, SUDOP, EKW, e-KRS, CEIDG, NSA and Eureka MF. How PL mobile proxies handle rate limits and datacenter ASN blocks. Traffic on our Polish mobile proxy ports shows a clear pattern: firms are not "scraping the web" in general - they hit specific Polish registry and government hosts. The largest volume goes to.

The June 4, 2026 refresh adds practical operating context for Scraping REGON GUS, RDF KRS, SUDOP, EKW and Polish Registries with Mobile Proxies: verify visible IP, DNS, ASN, protocol, rotation, session behavior, and target-platform response before using a proxy setup in production. This is a visible text update, not a date-only change, so the freshness signal is tied to reader-facing copy.

Polish mobile proxy for public registry data collection

Traffic on our Polish mobile proxy ports shows a clear pattern: firms are not "scraping the web" in general - they hit specific Polish registry and government hosts. The largest volume goes to the REGON GUS search (wyszukiwarkaregon.stat.gov.pl). Next: the Financial Documents Browser / RDF, SUDOP UOKiK, EKW land registers, Portal Rejestrów Sądowych / e-KRS, CEIDG, NSA judgments and Eureka MF tax rulings.

The data is public; rate limits, datacenter ASN blocks and PL geo still bite. This guide maps those eight sources and how to collect them through mobile proxies on Polish carriers - legally: respect limits, cache hard, keep normal session behaviour. No law-dodging tips.

  • Which hosts real client traffic hits (ordered by volume)
  • Why datacenter IPs fail on stat.gov.pl, ms.gov.pl and uokik.gov.pl
  • How to set limits, sticky sessions and cache per source type
  • When a PL mobile proxy cuts 429s and ASN bans

What clients actually scrape on PL proxies

This is not a list of "interesting APIs on paper". These are the hosts we see in business traffic: collections, KYC/AML, due diligence, counterparty monitoring, law firms, fintech, analytics. Pipelines usually start from an identifier (NIP, REGON, KRS, KW number), then add financial filings, state aid, land register, court entry, CEIDG, case law and MF interpretations.

Public data does not mean unlimited capacity at GUS, the Ministry of Justice, UOKiK or MF. A few hundred fast hits from a Hetzner/AWS IP get you 429, a captcha or a silent drop. A mobile IP on Play/Plus/Orange/T-Mobile in Poland looks like a domestic subscriber - and at a sane pace it lasts far longer.

REGON GUS search (wyszukiwarkaregon.stat.gov.pl)

Number one by client volume. The wyszukiwarka REGON on wyszukiwarkaregon.stat.gov.pl is the default start: NIP → REGON, entity status, address, PKD, start/end dates. Bulk counterparty checks and nightly CRM refreshes land here first.

Automation is fussy: sessions, form tokens, per-IP throttle. The classic mistake is "10 rps from one VPS in DE". Limits arrive fast. Better: sticky IP per batch (e.g. 50-200 NIPs), 0.3-1.5 rps with jitter, local cache by NIP/REGON with 24h-7d TTL. An active entity's REGON does not change every minute - full rescans without TTL only burn IPs.

PL geo matters: Polish mobile ASN traffic trips "foreign hosting" rules less often. After a job, confirm the exit with What Is My IP - mobile type, country PL.

Financial Documents Browser / RDF (rdf-przegladarka.ms.gov.pl)

Second by load: Przeglądarka Dokumentów Finansowych / Repozytorium Dokumentów Finansowych on rdf-przegladarka.ms.gov.pl. Financial statements, resolutions and KRS-linked attachments. Files are heavy; re-downloading the same PDFs every night looks aggressive and wastes bandwidth.

Practice: metadata / document list first, download only missing files (content hash in storage). Sticky session for one company's package. Do not rotate IP between "list" and "fetch PDF" - you drop context and farm 4xx. Keep rate limits lower than light REGON calls: PDFs are a different load class on MS infrastructure.

REGON GUS and RDF KRS scraping pipeline via PL mobile proxy

SUDOP UOKiK (api-sudop.uokik.gov.pl)

SUDOP (System Udostępniania Danych o Pomocy Publicznej) at UOKiK, including endpoints around api-sudop.uokik.gov.pl, covers state aid: who received support, on what basis, at what scale. Compliance, M&A due diligence and risk teams attach it to the counterparty card.

There is often an API layer - check the official channel and documented limits before you HTML-scrape. Keep a query budget, log 429/5xx, do not hide a flood behind rotation. PL mobile still helps around uokik.gov.pl UI/WAF, but a proper API client plus cache by entity id handles most volume without drama.

Electronic land registers EKW (przegladarka-ekw.ms.gov.pl)

The EKW browser on przegladarka-ekw.ms.gov.pl is Elektroniczne Księgi Wieczyste: KW content, sections, annotations. Traffic is bursty (real-estate DD, collections, banking). The UI is session- and pace-sensitive.

Rule: one sticky IP per KW or short batch of KW numbers. Do not fan out dozens of KW in parallel from one port. Cache with a short TTL (KW data is more "as of now" than REGON, but you still should not poll the same KW every 10 seconds in production). Captchas and limits show up faster on foreign datacenter than on Polish mobile.

Court registers portal / e-KRS (prs.ms.gov.pl)

Portal Rejestrów Sądowych and the KRS search on prs.ms.gov.pl (e-KRS) is the court entry: representation, capital, status, extracts. It often pairs with REGON (identify at GUS, confirm in KRS) and RDF (financial docs).

Bulk extracts burn sessions. Cache by KRS number, split light search from heavy extract downloads, backoff on 429. PL mobile proxy + sticky per entity cuts the "session tied to IP while you rotate every request" failure mode.

CEIDG and biznes.gov.pl

CEIDG (and paths via biznes.gov.pl) - sole traders: NIP, REGON, PKD, suspension, place of business. Still solid volume, though below REGON/RDF/SUDOP/EKW/PRS in our traffic. Forms and tokens break headless clients that skip step order.

Short bursts, pauses, sticky for the form session. Daily cache is usually enough for activity-status monitoring. Prefer official channels/API where they exist; HTML only for gaps.

NSA judgments (orzeczenia.nsa.gov.pl, CBOSA)

orzeczenia.nsa.gov.pl (administrative court case law, including the CBOSA context) is full-text judgments. Law and tax teams scrape for research, not for "10,000 NIP/h". Different profile: longer sessions, phrase search, fetching judgment bodies.

Stable session and low RPS matter more than aggressive rotation. PL mobile helps long research jobs that would otherwise exit from a foreign DC. Cache by judgment id; do not redownload what already sits in your corpus.

Tax interpretations Eureka MF (eureka.mf.gov.pl)

Eureka MF on eureka.mf.gov.pl is the Ministry of Finance tax ruling database. Lower volume than REGON, steady for tax/legal. Search + ruling text; same research pattern as NSA, not bulk NIP scoring.

Practice matches case law: calm pace, sticky, cache by signature/id, no rotation mid result pagination. Datacenter IPs get cut earlier when someone tries to "mirror the whole base over a weekend".

PL mobile proxy traffic to GUS, MS, UOKiK and MF hosts

Why datacenter loses and mobile PL holds the session

Shared across REGON, RDF, SUDOP, EKW, PRS, CEIDG, NSA and Eureka:

  • Hosting ASN - AWS, GCP, Azure, OVH, Hetzner land on "cloud" lists. Mass traffic is cheaper to cut than per-request behaviour analysis.
  • Non-PL geo - hits to stat.gov.pl / ms.gov.pl from Singapore raise risk scores.
  • Per-IP rate limits - 429 after N req/min; one VPS dies in a quarter hour.
  • Session bound to IP - rotating every URL breaks REGON, EKW and CEIDG forms more than it helps.

ProxyPoland 4G/5G mobile exits go through modems and Polish SIMs. Mobile ASN, CGNAT (many subscribers behind one public IP), PL geo. It is not a cloak of invisibility: 50 rps into REGON search still hits limits. At 0.2-2 rps per session with jitter and cache, the same port runs jobs for hours and days.

Shared good practice (no law-dodging)

Using public registries, case law and tax rulings for analytics and compliance is normal business practice in Poland. This is not an attack guide. Goal: less load on agencies, fewer bans on your side.

  1. Job order follows volume - REGON first (identity), then RDF/PRS/EKW as needed, SUDOP/CEIDG/NSA/Eureka as extra layers - not everything at once from one IP.
  2. Cache with TTL - NIP/REGON/KRS/KW/document id; fetched_at field.
  3. Sticky per entity or short batch; rotate between batches or after a limit error - not after every CSS GET.
  4. Backoff on 429 (30s → 2 min → 10 min), not instant IP swap and the same hammer.
  5. Official API where it exists (including SUDOP area) - HTML for gaps.
  6. Browser-like headers and cookies, real User-Agent, form step order preserved.
  7. Metrics - HTTP status, latency, captcha/WAF per host (REGON vs RDF vs EKW separately).

Protocol: HTTP(S) for light HTML; SOCKS5 when full Playwright/Puppeteer hits a heavy UI. Scale: one mobile port + cache = hundreds of REGON lookups/day; thousands plus parallel RDF PDFs = several ports with separate queues (separate RPS budgets for REGON vs PDF).

Bottom line

If your scraper targets the real Polish stack - wyszukiwarkaregon.stat.gov.pl, rdf-przegladarka.ms.gov.pl, api-sudop.uokik.gov.pl, przegladarka-ekw.ms.gov.pl, prs.ms.gov.pl, CEIDG/biznes.gov.pl, orzeczenia.nsa.gov.pl, eureka.mf.gov.pl - the bottleneck is IP reputation and pace, not the parser. Foreign datacenter burns fast. Mobile proxies on Polish carriers give PL geo, mobile ASN and CGNAT so limits and cache hold for days and weeks.

Order jobs by volume, cache identifiers, keep PDF and EKW RPS low. ProxyPoland: dedicated 4G/5G ports on Polish modems - trial, exit IP check, then production. Plans and trial · check IP

Before applying this article in production, verify the proxy protocol, visible IP, DNS route, ASN, target country, browser fingerprint, and rotation timing with the matching diagnostic tools. Treat the article as implementation guidance, then confirm the live setup against the current pricing and dashboard configuration.

FAQ

01Is scraping REGON, RDF, EKW, e-KRS, CEIDG, NSA, Eureka allowed?+

These are public information services. Firms have built KYC, scoring and research on them for years. That does not allow attacks, account bypass, flooding or breaking explicit bans. Stay on public endpoints, keep a sane pace, apply GDPR/RODO to downstream processing. Edge cases need legal review for your use case.

02Which host should the pipeline start with?+

Wyszukiwarka REGON GUS - highest volume and the natural NIP/REGON key. Then documents (RDF), state aid (SUDOP), KW (EKW), KRS entry (PRS), sole traders (CEIDG), research (NSA, Eureka) as the business needs them.

03Sticky or rotating on these hosts?+

Sticky for one entity or a short batch (several to fifteen minutes). Rotate between batches or after a limit error. On EKW and REGON forms, mid-path rotation is the most common cause of "weird" session errors.

04Why is "PL" residential sometimes worse than mobile?+

Paper geo is not always a Polish mobile carrier ASN. Mobile on a PL SIM is easy to verify: you see the operator and mobile type. On stat.gov.pl and ms.gov.pl that difference often shows up as job stability.