Recently I ran 20 Lighthouse passes with identical mobile settings on my own page and three pages that rank for "website audit". One page took a median 13.3 s to load its largest element (LCP) and still scored 71, while mine took 5.8 s and got only 69. So the scores alone rank a page 7.4 seconds slower higher (full measurement in my guide to benchmarking a website).
That's why a website audit report for a client opens with what to fix, on what basis, who'll do it and how you'll check the result. Below I build one in ten steps, and you can grab the templates and a fictional example report right now (the materials are in Polish).
What a website audit report for a client contains
A website audit report for a client is a document that turns review results into decisions and tasks with acceptance conditions. The decisions your client needs to approve sit at the top, with the scope, the findings and their evidence, the plan of work and the retest underneath.
Several people read it: the business owner wants decisions, the project lead settles scope, and the developer needs a description they can reproduce. I'd write one document in layers, with a summary for the decision-maker up top and the technical detail further down.
That's the split OWASP recommends for security testing reports, and I'm adapting it to a whole website report without calling it a standard.
| Part of the report | What the reader gets from it | What belongs there |
|---|---|---|
| Summary | A decision about further work | Goal, key findings, limitations and the proposed next step |
| Scope and method | How far the conclusions reach | URLs, templates, states, devices, sources, dates and exclusions |
| Findings | Material to do the work | Problem, evidence, impact, recommendation, owner and acceptance condition |
| Plan and retest | How the work is organised and closed | Order, dependencies, estimate, responsibility, retest result |
Every important finding answers five questions:
Your tool exports go in the appendices, and the main text points to a specific piece of evidence. W3C's template for accessibility evaluation reports has similar parts: the scope, the evaluation process and the results with recommendations.
The same outline works when you write up a completed audit.
Step 1: Settle the scope of the review before you write conclusions
The scope of the review is the list of URLs, page states and user tasks you actually checked. Writing "an audit of the whole site" doesn't say whether anyone checked the cart, the form after a failed submission, the mobile version or the logged-in dashboard.
Record the user tasks you tested: finding a product, placing an order, sending an enquiry, recovering an account. Then add the language, the consent state and the user role to each one.
Keep the inventory, the test plan and the testing actually done apart, because the sitemap only tells you what could be tested. Which URLs were actually opened and assessed after rendering is a separate record, and a timeout, a blocked request or a tool error goes in as a failed attempt.
For a larger site, describe how you picked the sample, because the homepage and two subpages don't represent every form, page template and state. WCAG-EM 2.0 covers sample selection and complete processes, and it warns that you can't claim WCAG conformance for a whole site based on a sampled subset alone.
Record the number of URLs, assessed templates and checked scenarios separately, because one URL can hold several important states.
When the total number of pages is unknown, write "total count not established". A coverage percentage from an unknown denominator is a number pulled out of thin air.
Put your exclusions next to the summary. Without access to payments, you can't confirm the full purchase, so "the shop works correctly" is off the table.
Step 2: Describe the evidence so someone else can check it
Evidence is a record that lets someone else reproduce the same observation. A finding gets a stable identifier, such as A11Y-01, kept apart from the identifiers of its occurrences. The same header link can sit on many pages and need one change to a shared component.
With the evidence, record:
- the URL or the path to reach it;
- the date with a time zone;
- the environment and the page state;
- the tool name and version;
- a description of the result, plus the unit and settings for measurements.
A screenshot helps you find the element, but it's the accessibility tree, the HTTP headers or a request log that explain the observation.
Those fields protect you from a number that only looks like a measurement. I got caught by one myself when GA4 recently showed me 70 sessions on cometweb.io over 28 days.
The source/medium report broke them down like this:
In reality, that's about 1 session a day.
Without a source, a period or a channel split, "70 sessions" would've gone into a report as a fact.
Keep the observation, the conclusion about impact, the hypothesis about the cause and the recommendation separate. A link missing a programmatically determinable name is an observation.
Someone using a screen reader may find it hard to tell where the link goes. Saying so is a justified description of the barrier, according to W3C's guidance on SC 4.1.2 Name, Role, Value. A drop in sales by a given percentage needs different data.
You check focus and keyboard activation separately from the name, and before you add aria-label everywhere, check whether proper link text is enough.
Until the cause is confirmed in the code, mark it as a hypothesis, because similar warnings can come from different components. Group them only once you've confirmed a shared cause, or with a note that the grouping is a working assumption.
Step 3: Fill in a finding card, using a worked example
A finding card is one table where the problem and the evidence sit next to the owner and the acceptance condition. You can download the empty finding card (in Polish).
Fictional scenario. Every URL, evidence identifier and number in the card is illustrative. Nobody tested the shop, and the example is for learning only: it shows neither client measurements nor CometWeb results.
| Field | Example entry |
|---|---|
| Identifier and title | A11Y-01: "My account" link with no programmatically determinable name |
| Scope of observation | In this scenario: 6 of 6 checked pages that contain this header variant; other states not established |
| Evidence | Fictional record E-01: role link, empty name, target /account; element located in the header |
| Impact | Someone using a screen reader may find it hard to identify where the link goes; no revenue impact was estimated |
| Cause | Hypothesis: the shared component contains only an icon; needs confirming in the code |
| Recommendation | Give the link a descriptive name, preferring proper link text, and check every variant of the component |
| Priority and owner | Proposed P2; the developer responsible for the component; scope approved by the project lead |
| Acceptance condition | In every variant checked, the accessibility tree shows the correct name and role; focus, Enter and the target page were checked separately |
| Status | Scope and implementation to be confirmed; no retest performed yet |
With a card like this, you can commission the work without explaining it in a meeting.
Step 4: Separate priority, severity and confidence
Severity describes the effect of a problem, priority sets the order of work in a project, and confidence says what's actually been confirmed.
An uncertain signal of a serious threat needs quick verification, even while the evidence is still missing. The priority split below is my own proposal for organising the work.
| Priority | When to consider it | Decision for the team |
|---|---|---|
| P1: urgent | A confirmed blocker in an important process, or a significant threat that needs a response | Name the person responsible, limit the impact and schedule an urgent test |
| P2: next scope of work | A significant, confirmed barrier whose removal clearly matters | Plan the fix and the sign-off |
| P3: scheduled | An improvement with less impact, or one that depends on earlier work | Set a date once the dependencies are resolved |
| P4: watch | A signal without enough context, or a low priority for improvement | State what data will decide the next step, and when |
An error in a shared template doesn't go to P1 automatically. A minor problem across many pages can be less urgent than a broken form on one important page. A small number of users, though, doesn't remove an accessibility barrier.
Ease of fixing helps you order the work within a priority. Add dependencies to your tasks: link fixes may wait for a decision on the target URLs, and a cache change for the rules on the cart.
Step 5: Present results from different areas without mixing meanings
Each audit area has its own data source and its own limits on what you can conclude, so your report shows them separately.
Performance: the measurement source before the number
For Core Web Vitals (LCP, INP, CLS), say where the result comes from: a lab test, CrUX or your own monitoring of real visits. Add the URL or origin, the device type and the period.
According to web.dev and the CrUX API documentation, data for an origin (such as https://www.example.com) covers all its pages together and doesn't measure each individual page.
Lighthouse gives different results on consecutive runs, as Chrome for Developers explains in its performance scoring docs. In your report, show the number of runs, the test conditions and the full spread.
On cometweb.io/website-audit, Lighthouse (mobile) LCP moved by 886 ms between the fastest and the slowest of five runs on one morning in September 2026.
INP measures interactions, and TBT from a load test is only a supporting signal.
Good Core Web Vitals values are LCP up to 2.5 s, INP up to 200 ms and CLS up to 0.1, at the 75th percentile of visits according to web.dev. A median of a few lab runs is still lab data, while a field p75 comes from real visits.
Leave aggregate scores to tools that document how they're built: an average of performance, SEO or security points sold as a "site health percentage" mixes different scales. If you quote a tool's score, give its name and methodology, along with its scope and data gaps. A single blocker in a process can need action even when the rating is high.
SEO: page configuration and Google's decision go in separate places
A canonical in the HTML, the URL Google selects and visibility in search results are three different observations. Document the page source and what URL Inspection tells you separately, with the last crawl date. The live URL test in Search Console describes the current version, and a separate view shows the version in the index.
Don't attach a redirect to every canonical recommendation. It suits a URL you're retiring, and for page variants you still need, the decision may be different. According to Google Search Central on canonical URLs, consistent signals help indicate the preferred URL, but the choice is Google's.
Valid structured data gives a page a chance at a rich result, and Google makes that call too (general structured data guidelines). Name the tool, the type checked and the scope of the check, because "schema: passed" tells your client nothing.
Accessibility: an automated scan checks some of the criteria
For each accessibility finding in your report, give the criterion, the WCAG version, the evaluation method and the state tested. Supplement your automated results with manual testing and assistive technologies, because according to W3C WAI, no tool checks every aspect of accessibility.
List what you checked and what you skipped, because otherwise the reader will take the absence of a warning as conformance.
Don't present scanner points as "96% WCAG compliance". W3C's explanation of conformance defines conformance requirements, and WCAG has no percentage like that. Legal obligations need their own assessment.
I cover the line between automated and manual evaluation in my post on a WCAG and EAA accessibility audit.
Security and CO₂e: a conclusion as wide as the review
A check of HTTP headers and TLS is a slice of a security test. Record whether you analysed authentication and roles, as well as APIs and logged-in data.
Telling anyone "the site is secure" after a limited review is a promise with nothing behind it. OWASP stresses that a test is a point-in-time assessment and can miss some issues.
If your report includes CO₂e, call the result an estimate and give the model, the version, the inputs and the boundaries of the calculation (Green Web Foundation, CO2.js models). A transfer-based model estimates emissions from the data sent and doesn't read the energy a particular phone used. So compare before and after with the same method.
Step 6: Write the summary for the client
Your summary connects the goal, the main problem, the proposed action and the limits of what's known. A short opening section is usually enough.
Fictional example:
The goal of the review was to prepare the shop for fixes ahead of the next campaign. In this scenario, six public pages and selected interactions were assessed. We recommend first giving the "My account" link in the shared header a name, and resolving the conflicting canonical signals on two pages. The team needs to approve the target URLs before the SEO changes are deployed. Real payment, the logged-in dashboard and the full WCAG scope were not tested. Without INP field data, it is not possible to assess how the site responds for real users. The client should approve the scope of the first two tasks, name the person responsible for SEO decisions and arrange access for the retest.
Leave any unfounded forecast of sales growth out of your summary. According to Google on page experience, a good Core Web Vitals result won't secure high rankings, and a technical improvement, a traffic trend and a business result each need their own evidence.
Step 7: Assign owners, cost and dependencies in the implementation plan
The implementation plan gives every task one owner for delivery and one person who'll accept the result: a developer fixes the component, QA runs the retest. The client decides on scope and leaves judging the code to the technical team. If you're only proposing the responsibility, label it that way in the plan.
The suggested owner usually falls into one of five roles, which the organisation assigns in the end:
| Role | What they own |
|---|---|
| Developer | code, routing, headers, performance and components |
| SEO | indexing, canonicals, sitemap, structured data and internal linking |
| Content | titles, definitions, missing answers and text quality |
| Design and accessibility | contrast, focus, interaction order and messages |
| Owner or marketing | business priority, scope approval and deadline |
Price diagnosis, implementation, testing, deployment and production checks as separate lines, with your assumptions. When access to the code or an external vendor is missing, write "to be estimated after diagnosis" where a falsely precise number of hours would go.
Agree with your client what the engagement covers (the review, consultation or implementation) and how many retest rounds it includes. A recommendation is a proposal, and ordering the work is a separate decision, just like alternative solutions and extra work.
Keep the state recorded in your report separate from your live task list. Keep the version your client received, link tasks back to the finding identifiers, and add the "after" image next to the earlier evidence.
Step 8: Retest the fix before you close it
A retest repeats the same scenario after the fix and checks it against an acceptance condition written before the work began. The condition states what's being checked, the expected result and the method, and "the score should be green" is far too vague for that.
After deployment, reproduce the scenario and check the important variants and regressions. In your report, record the work status and the verification result as separate fields, for example on the retest and sign-off card (in Polish).
For a risk acceptance, record who decided and why, plus a date for the next review, and the technical result stays negative.
For performance, you compare the same configurations and the same statistic, and flag a change of tools, review scope or device as a limit on the comparison. A gap between mobile and desktop data comes from the devices themselves, so describe it separately from the effect of the change.
According to the CrUX API documentation, data updates daily over a rolling 28-day window, so after a fix it mixes visits from before and after the change. Record the data collection period in your report and leave out any promise of a date on which the result will definitely improve.
Doing this by hand takes a lot of tracking. In CometWeb Insight, our website audit app, results from every module go into a prioritised report and a task queue with statuses. A re-scan with the same scope and the same methodology shows what's been fixed, what's still open and what's new.
A "done" status is the cue for a re-scan, and for automatically detected findings the re-scan itself is the evidence. You check manual acceptance conditions as described above.
Step 9: Hand over the report so it's readable and secure
A readable report has descriptive headings, understandable link text and statuses written as text, readable without colour too. Each screenshot gets a description of the barrier or observation (W3C WAI on writing for accessibility).
HTML makes navigation easier, and a PDF works as a frozen handover version.
In either, you check the structure, the reading order, how the links work and whether it's readable on a small screen. A set of images saved as a PDF fails.
Send your confidential report through a channel with access control, with agreed recipients, a retention period and a way to revoke access. Strip a public example of personal data, tokens, session identifiers, admin URLs and details of an unfixed vulnerability. You store the original evidence separately.
noindex doesn't protect a confidential report, because anyone who has the link can still open the page. Google distinguishes excluding a page from its index from protecting content against unauthorised access, and a hard-to-guess URL tells you nothing about who the recipient is.
Step 10: End the client meeting with decisions
A client meeting should end with a list of decisions, so send your summary and main findings beforehand.
During the meeting, show your most important evidence, explain the extent of the impact and propose an order. Skip reading every warning out in turn.
At the end, record the agreed tasks with owners and acceptance conditions, the postponed work and the open questions. If the client needs more to decide, work out what information that is. After the meeting, update the register and link it to the report version.
Three takeaways on audit reports
- A conclusion reaches as far as the review. An export with no priorities, or a report with no scope and no date, leaves your client to translate the data. An automated accessibility audit isn't a certification.
- Keep the observation and the cause apart from the recommendation. Combine occurrences with a confirmed shared cause into one task.
- A retest closes the work. Without an owner and an acceptance condition for every finding, your team doesn't know when it's done.
Any tool will produce the list of problems. For the rest, start from the report template (in Polish), where empty fields are marked as not established so nobody mistakes them for positive results.
The sample Insight report shows how results are presented, and the CometWeb methodology explains the sources, missing data and priorities. If you'd like every area of website quality in one view, with a task queue and re-scans, create your CometWeb Insight account.
Frequently asked questions about website audit reports
Is a Lighthouse export enough as a client report?
It can be an appendix or the result of a narrowly defined engagement. A broader audit needs interpretation, scope, evidence, priorities and a verification plan added. The result of a single run describes the conditions of that run, not every visit.
How many pages should an audit report have?
As many as it takes to communicate the findings and support a decision. A short document can be complete for a small scope. A large site usually needs extensive appendices. The page count does not replace a description of the sample, the states and the methods.
Should every warning be a separate task?
No. Combine occurrences with a confirmed shared cause, and keep the list of locations and variants for the retest. A different cause, owner or verification method can justify a separate task.
How do you describe missing data?
State what is missing, why, and which conclusion is therefore unavailable. Distinguish between "not tested", "test failed", "no field data" and "not applicable". A value of zero is a result, not a stand-in for a measurement that was never taken.
When can you say a problem has been fixed?
When a retest result confirms the previously agreed condition within the stated scope. A commit, a deployment, a closed ticket or a decision to accept the risk is not enough on its own. In every other case, record the current work status and the verification that is still missing.
Is a score enough as a summary?
A score can orient the reader, but it should not drive the decision. A number without scope, source and date hides gaps in coverage and mixes measurement with recommendation. Lead the report with findings and evidence, and treat the score as a label, not as the content.
Sources and scope of this guide
This is a guide to preparing a technical report. The examples are fictional; they are not an audit, a penetration test, a conformance statement or a quote for any particular site. The priority split, the card templates and the sign-off rules are a proposal for organising the work, with no status as a new technical standard.
- OWASP WSTG: Reporting. Audiences, scope, limitations, version control and protecting the report; the guide now lives on wstg.owasp.org.
- W3C: WCAG Evaluation Methodology 2.0. Sampling, complete processes, the scope of claims and the limits of scoring. Group Note of 23 July 2026.
- Chrome for Developers: CrUX API. URL/origin scope, device, daily updates and the rolling 28-day window.
- web.dev: Web Vitals. LCP, INP and CLS, the 75th percentile and the good thresholds.
- Google Search Central: Canonical URLs. A canonical and a redirect are different decisions; signals do not guarantee Google's choice.
- Google Search Console: URL Inspection tool. The indexed version, the live URL test and the limits of interpretation.
- W3C WAI: Selecting Web Accessibility Evaluation Tools. Tools support evaluation but cannot check every aspect of accessibility.
- W3C WAI: Understanding SC 4.1.2: Name, Role, Value. Programmatically determinable name and role; the link example.
- W3C WAI: Understanding Conformance. WCAG conformance requirements, full pages and complete processes.
- Green Web Foundation: CO2.js models. Emission estimates, model version and system boundaries.
- Google Search Central: Control what you share. Noindex does not block direct access; protecting confidential content.
- Google Search Central: Understanding page experience. A good CWV result does not guarantee top rankings.