CometWebPakiet badania

Internal link audit of cometweb.io

Maciej Zmitrukiewicz · Pakiet badania

This is the material behind the article Internal link audit: find orphan pages and weak links and its Polish version, Linkowanie wewnętrzne: audyt i strony osierocone. In September 2026 I crawled my own site, cometweb.io, to count its internal links, orphan pages and click depth. No other site was crawled. A Polish summary is at the end.

What I looked at

The crawler started at the English homepage (/) and the Polish one (/pl) and followed every <a href> in the HTML the server sends. It sent one request at a time with a 0.6 s pause after each, read robots.txt first and obeyed it, and identified itself as CometWeb-LinkAudit/1.0. It kept no cookies and sent no Accept-Language, so the site's locale redirect couldn't carry over from one request to the next. It recorded any redirect hop by hop instead of following it silently, although none came up.

For every link it saved the source page, the href as written, where it resolves, the anchor text, rel, the part of the page the link sits in (header, footer, body and so on) and the target's status. Sitemap URLs the crawl never reached were requested at the end so their status is known, but their own links weren't followed, so they couldn't make themselves reachable.

A separate browser check loaded eight pages (both homepages, both blog indexes, /tools, /glossary and both versions of the article) and compared the internal link targets in the rendered page with those in the server HTML.

How I counted

The depth and inbound-link figures cover indexable pages only: no query string, no noindex, a canonical pointing at the page itself and outside /research/. That came to exactly the 273 sitemap URLs the crawl reached.

Results

What I countedResult
Sitemap URLs281
HTML pages reached from the two homepages337 (174 EN, 160 PL, 3 DE)
of which indexable273 (135 EN, 135 PL, 3 DE)
of which parameter versions of another page37, all with a canonical to the clean URL
of which noindex product pages10
of which /research/ pages17
All links on crawled pages20,252
Internal links on crawled pages15,850 (11,516 unique source and target pairs)
Internal links per indexable pagemedian 47, range 37–93
Internal links in footer / header / top-level nav / sidebar10,596 / 2,131 / 0 / 200
Internal links in body text / in-page nav in the body2,241 / 682
Orphan pages (in the sitemap, never reached)8: four English blog posts and their Polish versions, all returning 200, indexable and with a self-canonical
Links to redirects0
Links to 4xx or 5xx0
Internal links with nofollow0
<a> without href0 in the server HTML and 0 in the rendered pages I checked
Click depth of indexable pagesEnglish from /: 1 page at 0, 41 at 1, 93 at 2. Polish from /pl: the same. Nothing deeper than 2
Indexable pages linked from fewer than 2 pages outside header, nav and footer24, including 8 with none: legal pages, /bot and /insight/mcp in both languages
Generic anchors9 links on 5 pages
Same anchor text pointing to different pages, outside the template30 anchor texts
Parameter links44 links from 30 pages to 37 URLs: 40 to /contact?topic=… and 4 to /pl/insight/…?source_page=…
Server HTML vs rendered page (8 pages)the same internal link targets on every page

Two things I fixed before publishing

My first crawl treated every <nav> inside the body, such as breadcrumbs, the table of contents and related posts, the same as the site header. That moved contextual links into the template bucket and made inbound links look thinner than they are. I separated in-page navigation from the header and ran the whole crawl again, and only that second run is published here.

I also compressed links.csv to links.csv.gz after the run. analyse.py reads either form and gives the same summary.json.

What this data can't tell you

Files

To run it again, use Python 3.10 or later: python3 crawl.py, then python3 analyse.py, and optionally npm i playwright && node render-check.mjs. Crawl only sites you own or are allowed to audit.

Back to the article

Streszczenie po polsku

We wrześniu 2026 przecrawlowałem własną stronę, cometweb.io, zaczynając od strony głównej EN i PL. Brałem tylko linki <a href> z HTML-u, który zwraca serwer, wysyłałem jedno żądanie naraz z przerwą 0,6 s i trzymałem się robots.txt.

Crawl dotarł do 337 stron HTML, z czego 273 są indeksowalne, i znalazł 15 850 linków wewnętrznych. Osieroconych stron było 8: to wpisy blogowe w sitemapie, do których nie prowadzi żaden link. Linków do przekierowań i błędów było 0, wewnętrznych nofollow też 0, a żadna indeksowalna strona nie leży dalej niż 2 kliknięcia od strony głównej swojej wersji językowej. Do tego 9 ogólnych anchorów i 44 linki do adresów z parametrami.

Definicje, ograniczenia i dwie poprawki metody są opisane wyżej, a skrypty pozwalają odtworzyć każdą liczbę. Wróć do artykułu