Internal link audit of cometweb.io
This is the material behind the article Internal link audit: find orphan pages and weak links and its Polish version, Linkowanie wewnętrzne: audyt i strony osierocone. In September 2026 I crawled my own site, cometweb.io, to count its internal links, orphan pages and click depth. No other site was crawled. A Polish summary is at the end.
What I looked at
The crawler started at the English homepage (/) and the Polish one (/pl) and followed every <a href> in the HTML the server sends. It sent one request at a time with a 0.6 s pause after each, read robots.txt first and obeyed it, and identified itself as CometWeb-LinkAudit/1.0. It kept no cookies and sent no Accept-Language, so the site's locale redirect couldn't carry over from one request to the next. It recorded any redirect hop by hop instead of following it silently, although none came up.
For every link it saved the source page, the href as written, where it resolves, the anchor text, rel, the part of the page the link sits in (header, footer, body and so on) and the target's status. Sitemap URLs the crawl never reached were requested at the end so their status is known, but their own links weren't followed, so they couldn't make themselves reachable.
A separate browser check loaded eight pages (both homepages, both blog indexes, /tools, /glossary and both versions of the article) and compared the internal link targets in the rendered page with those in the server HTML.
How I counted
- Inbound links are the number of different pages that link to a page, without self-links. A link to a redirect counts for its final URL. I counted them three ways: all links, links outside the header, top-level nav, footer and sidebar, and links in the body text only.
- An orphan page is a URL in
sitemap.xmlthat the crawl from the two homepages never reached. - Click depth is the shortest path in clicks from the page's own homepage:
/for English and German,/plfor Polish. - A generic anchor is an exact match against a short list in
analyse.py: Google's own examples ("click here", "read more", "website", "article"), close variants such as "learn more" and "view all", and Polish equivalents such as „czytaj więcej” and „dowiedz się więcej”. "Read more about X" doesn't count. - A parameter link points to an internal page URL with a query string.
The depth and inbound-link figures cover indexable pages only: no query string, no noindex, a canonical pointing at the page itself and outside /research/. That came to exactly the 273 sitemap URLs the crawl reached.
Results
| What I counted | Result |
|---|---|
| Sitemap URLs | 281 |
| HTML pages reached from the two homepages | 337 (174 EN, 160 PL, 3 DE) |
| of which indexable | 273 (135 EN, 135 PL, 3 DE) |
| of which parameter versions of another page | 37, all with a canonical to the clean URL |
of which noindex product pages | 10 |
of which /research/ pages | 17 |
| All links on crawled pages | 20,252 |
| Internal links on crawled pages | 15,850 (11,516 unique source and target pairs) |
| Internal links per indexable page | median 47, range 37–93 |
| Internal links in footer / header / top-level nav / sidebar | 10,596 / 2,131 / 0 / 200 |
| Internal links in body text / in-page nav in the body | 2,241 / 682 |
| Orphan pages (in the sitemap, never reached) | 8: four English blog posts and their Polish versions, all returning 200, indexable and with a self-canonical |
| Links to redirects | 0 |
| Links to 4xx or 5xx | 0 |
Internal links with nofollow | 0 |
<a> without href | 0 in the server HTML and 0 in the rendered pages I checked |
| Click depth of indexable pages | English from /: 1 page at 0, 41 at 1, 93 at 2. Polish from /pl: the same. Nothing deeper than 2 |
| Indexable pages linked from fewer than 2 pages outside header, nav and footer | 24, including 8 with none: legal pages, /bot and /insight/mcp in both languages |
| Generic anchors | 9 links on 5 pages |
| Same anchor text pointing to different pages, outside the template | 30 anchor texts |
| Parameter links | 44 links from 30 pages to 37 URLs: 40 to /contact?topic=… and 4 to /pl/insight/…?source_page=… |
| Server HTML vs rendered page (8 pages) | the same internal link targets on every page |
Two things I fixed before publishing
My first crawl treated every <nav> inside the body, such as breadcrumbs, the table of contents and related posts, the same as the site header. That moved contextual links into the template bucket and made inbound links look thinner than they are. I separated in-page navigation from the header and ran the whole crawl again, and only that second run is published here.
I also compressed links.csv to links.csv.gz after the run. analyse.py reads either form and gives the same summary.json.
What this data can't tell you
- How Google sees the site. This is my crawler, not Googlebot, and I built the orphan list from the sitemap only, so a page that is neither linked nor in the sitemap stays invisible here.
- How much a link weighs. Inbound links are counts of linking pages.
- Links that appear only after you open a menu or press a button. The browser check didn't click anything.
- How the site looks later. I deploy several times a week, so running the scripts again will give different numbers.
- The three German pages are counted, but I didn't analyse them as a separate language.
Files
summary.json: every number quoted in the articlepage-metrics.csv: inbound links, click depth and outgoing internal links for each pageorphans.csv: the 8 orphan pagesgeneric-anchors.csv: the 9 generic anchorsredirect-links.csv: links to redirects (empty, there were none)render-check.json: server HTML compared with the rendered pagelinks.csv.gz: every link the crawl recordedpages.csv,resolved-urls.csv,fetches.csv,sitemap-urls.txt,crawl-run.json: raw crawl data and the run settingssources.json: the outside sources I relied on, such as Google's link guidelinescrawl.py,analyse.py,render-check.mjs: the crawler, the analysis and the browser check
To run it again, use Python 3.10 or later: python3 crawl.py, then python3 analyse.py, and optionally npm i playwright && node render-check.mjs. Crawl only sites you own or are allowed to audit.
Streszczenie po polsku
We wrześniu 2026 przecrawlowałem własną stronę, cometweb.io, zaczynając od strony głównej EN i PL. Brałem tylko linki <a href> z HTML-u, który zwraca serwer, wysyłałem jedno żądanie naraz z przerwą 0,6 s i trzymałem się robots.txt.
Crawl dotarł do 337 stron HTML, z czego 273 są indeksowalne, i znalazł 15 850 linków wewnętrznych. Osieroconych stron było 8: to wpisy blogowe w sitemapie, do których nie prowadzi żaden link. Linków do przekierowań i błędów było 0, wewnętrznych nofollow też 0, a żadna indeksowalna strona nie leży dalej niż 2 kliknięcia od strony głównej swojej wersji językowej. Do tego 9 ogólnych anchorów i 44 linki do adresów z parametrami.
Definicje, ograniczenia i dwie poprawki metody są opisane wyżej, a skrypty pozwalają odtworzyć każdą liczbę. Wróć do artykułu