A Sitebulb audit is a structured crawl that turns technical SEO data into an action plan. This guide covers 12 steps, from project setup and crawler selection to manual validation and a follow-up crawl.
Publication date:
Sitebulb Audit Setup at a Glance
| Audit requirement | Recommended Sitebulb setting |
|---|---|
| Standard technical SEO crawl | HTML Crawler for speed, or Chrome Crawler when JavaScript affects page content |
| JavaScript SEO | Chrome Crawler |
| Performance and mobile analysis | Chrome Crawler with Performance & Mobile Friendly enabled |
| Orphan page discovery | Connect XML sitemaps, Google Analytics and Google Search Console |
| Internal linking analysis | Enable Search Engine Optimisation and Link Analysis |
| Image, JavaScript and CSS checks | Enable Page Resources |
| Large website | Run a sample crawl first and enable only the required reports |
| Client reporting | Use the PDF Report, URL exports or Google Sheets exports |
Sitebulb separates a project from an audit. A project stores the crawl configuration. An audit is one completed crawl and the reports produced from it.
1. Define the Audit Scope
Decide what you need to investigate before opening Sitebulb. A full technical SEO audit may include:
- Crawlability and indexability
- Robots.txt and XML sitemaps
- Canonical tags
- HTTP status codes
- Internal links and redirects
- Orphan pages
- Title tags, meta descriptions and headings
- Duplicate or similar content
- JavaScript-rendered content
- Page resources
- Performance and mobile friendliness
- Structured data and hreflang, where relevant
Do not enable every report by default. Each data set increases crawl time and computer resource usage. Performance, accessibility and similar-content analysis can require more processing than a standard crawl.
2. Create a Sitebulb Project
In Sitebulb:
- Create a new project.
- Enter the canonical homepage or the relevant subfolder as the Start URL.
- Give the project a clear name.
- Select the crawler type.
- Click Save and Continue.
- Review the audit settings before starting the crawl.
Use the preferred URL version. For example, if the site uses `start with that version rather than an HTTP or non-www variation.
If the site is protected by authentication, staging restrictions or an allowlist, set up access before starting the audit. Otherwise, Sitebulb may crawl no URLs or return incomplete results.
3. Choose the Right Crawler
The crawler determines which version of each page Sitebulb can inspect.
Which sites should use the HTML Crawler?
The HTML Crawler extracts information from the server response. It is usually faster than the Chrome Crawler and suits websites where important content and links are present in the initial HTML.
Use it when:
- The site is mostly server-rendered
- You need to crawl a large site quickly
- JavaScript does not create important content or links
- You are checking status codes, metadata and crawl paths
When should you use the Chrome Crawler?
The Chrome Crawler uses headless Chrome to render pages. Choose it when JavaScript creates or changes important content, navigation, internal links, canonical tags or meta robots directives.
Use the Chrome Crawler when:
- The site uses React, Vue, Angular or another JavaScript framework
- Content appears only after rendering
- JavaScript creates internal links
- You need performance or mobile-friendly data
- You want to compare response HTML with rendered HTML
If you are unsure, run a small sample crawl with the Chrome Crawler and inspect the Response vs. Render report. It shows whether JavaScript creates, changes or removes important SEO elements.
4. Configure the Audit Data
The settings depend on the audit scope. For a standard technical SEO audit, enable the following where they apply:
- Search Engine Optimisation
- Link Analysis
- Page Resources
- Performance & Mobile Friendly
The Search Engine Optimisation report collects core SEO information, including internal links and indexability signals. Page Resources adds checks for JavaScript, CSS, images, video and audio. Performance and Mobile Friendly requires the Chrome Crawler.
For a smaller or faster crawl, start with:
- Search Engine Optimisation
- Link Analysis
- XML Sitemaps
Add performance, accessibility, spelling or duplicate-content reports only when they fit the audit.
Configure Performance Data Separately
Sitebulb collects performance data through headless Chrome. Its default Web Vitals sampling recommendation is 10%, although performance-related Hints are still checked across crawled pages. You can also select a mobile or desktop device.
Use mobile as the primary device when mobile usability and performance are part of the audit scope.
5. Add Every Important Crawl Source
A crawl that follows internal links alone can miss useful URLs. Add other sources to create a fuller URL inventory.
Connect XML Sitemaps
Enable XML sitemap crawling and add known sitemap URLs, including:
/sitemap.xml- Sitemap index files
- Product sitemaps
- Article or blog sitemaps
- Image or video sitemaps, where relevant
Sitebulb can find sitemap URLs through robots.txt or Google Search Console. You can also add them manually.
Connect Google Analytics
Connect Google Analytics and enable the option to extract and crawl URLs found in the property. This can reveal URLs that receive analytics data but are not reachable through internal links.
Connect Google Search Console
Connect Google Search Console and allow Sitebulb to extract and crawl URLs found there. Search Console can reveal indexed or previously visible URLs that the current internal linking structure no longer exposes.
Upload a URL List
Upload a URL list when you need to audit a defined group, such as:
- Key landing pages
- A product catalogue
- A migration URL set
- URLs from a previous crawl
- Pages supplied by a client or development team
6. Run the Crawl and Check the Setup
Start the audit and wait for Sitebulb to finish crawling.
Before reviewing the findings, check:
- The number of URLs crawled
- Whether the homepage returned the expected status code
- Whether Sitebulb crawled the correct subdomain
- Whether robots.txt blocked important areas
- Whether XML sitemap URLs were accepted
- Whether Google Analytics and Search Console connected successfully
- Whether rendered pages contain the expected content
- Whether the crawl stopped early
A surprisingly low URL count often points to a configuration, robots.txt, authentication, redirect or internal-linking problem. It does not necessarily mean the site is small.
7. Review the Audit Overview and Hints
When the crawl finishes, Sitebulb provides an Audit Overview with crawl data, scores, reports and triggered Hints. The audit score is based on the type, number and priority of Hints found during the crawl.
Use the score to understand the general state of the crawl, but do not treat it as the final diagnosis. A site can have a high score and still contain one serious issue on an important commercial page.
Sitebulb groups Hints into five priority levels:
- Critical
- High
- Medium
- Low
- Insight
Open each relevant Hint, inspect the affected URLs and review the underlying evidence before recommending a fix. Sitebulb often shows the issue directly in the page or crawl data.
8. Audit Indexability and Crawlability
Open the Indexability report and review:
- Indexable URLs
- Noindex URLs
- Nofollow URLs
- URLs blocked by robots.txt
- Canonical status
- Meta robots directives
- HTTP status codes
- Robots.txt rules
- Redirected and error URLs
The main question is whether the pages that should appear in search engines are crawlable, indexable and represented by the correct canonical URLs.
Review important page types separately. Category pages, product pages, service pages and editorial content may use different indexability rules by design.
Review Canonical Tags
Sitebulb groups pages according to whether their canonical points to:
- Themselves
- Another internal URL
- An external URL
- No URL because the canonical is missing
Investigate canonicals that point to:
- Redirecting URLs
- 404 pages
- 5xx pages
- Noindex pages
- Disallowed URLs
- Another canonicalized URL
- An incorrect HTTP version
Sitebulb includes Hints for several of these canonical problems.
Compare Response and Rendered HTML
For JavaScript websites, use the Response vs. Render report to compare the original HTML response with the rendered page. Check for changes to:
- Title tags
- Meta descriptions
- Meta robots
- Canonical tags
- Internal links
- External links
A page can look correct in a browser while providing incomplete or conflicting signals before JavaScript runs.
9. Audit Internal Links and Orphan Pages
Open the Links report to review:
- Broken internal links
- Redirected internal links
- Links to HTTP URLs
- Pages with few internal links
- Deep pages
- Internal link distribution
- Orphan URLs
- External link problems
Fix broken internal links first. Then check whether important pages have enough relevant links from the sections of the site where users and search engines would expect to find them.
Find Orphan Pages
Sitebulb defines an orphan page as a URL found through another source but not discovered during the internal crawl. You can identify these pages by comparing the crawl with XML sitemaps, Google Analytics and Google Search Console.
Prioritise orphan pages that:
- Receive organic traffic
- Generate conversions
- Appear in XML sitemaps
- Have backlinks
- Rank for valuable queries
- Represent current products or services
An orphan URL is not automatically a problem. Some pages should remain outside the main navigation. Others became disconnected by mistake.
10. Review On-Page SEO and Duplicate Content
Use the On Page report to check:
- Missing title tags
- Duplicate title tags
- Multiple title tags
- Missing or duplicate meta descriptions
- Missing or multiple H1 headings
- Overly long or short elements
- Thin pages
- Duplicate or similar content
- Missing image alt text, when Page Resources is enabled
Start with patterns affecting many URLs. A template-level title problem across 5,000 product pages usually matters more than one imperfect title on a low-value blog post.
Sitebulb's On Page report includes title identification data and Hints for missing or multiple title elements.
11. Review Performance, Mobile and Page Resources
If you enabled Performance & Mobile Friendly, review:
- Core Web Vitals data
- Slow page templates
- Large JavaScript and CSS files
- Render-blocking resources
- Image sizes and formats
- Mobile usability problems
- Performance opportunities identified by Sitebulb
Look for groups of pages affected by the same template or asset. A single laboratory performance result does not show how every user experiences the site. Compare important findings with real-user data where available.
12. Prioritise Findings by Business Impact
Do not send every Sitebulb Hint to a developer. Turn the findings into a prioritised list.
Priority 1: Fix Issues That Can Block Visibility
Examples include:
- Important pages blocked by robots.txt
- Accidental noindex directives
- Incorrect canonical tags across key templates
- Server errors on commercial pages
- Broken domain or HTTPS redirects
- JavaScript preventing important content or links from rendering
Priority 2: Fix Widespread Structural Problems
Examples include:
- Broken internal links
- Orphaned commercial pages
- Large groups of duplicate titles
- Weak internal linking to important sections
- Incorrect XML sitemap URLs
- Duplicate or indexable parameter pages
Priority 3: Improve Page-Level Quality
Examples include:
- Missing meta descriptions
- Minor title-length issues
- Isolated heading problems
- Low-impact image alt text issues
- Small performance opportunities on low-value pages
For every recommendation, record:
- The affected URL pattern
- The issue
- The evidence
- The business impact
- The recommended fix
- The validation method
13. Export the Findings and Re-Audit
Sitebulb provides URL Explorer, Link Explorer, bulk exports, PDF reports and Google Sheets integrations for reviewing and sharing crawl data.
A useful deliverable contains:
- Executive summary
- Critical technical issues
- Indexability and canonical findings
- Internal linking and orphan-page findings
- On-page issues
- Performance and JavaScript findings
- Prioritised recommendations
- URL examples and implementation notes
- Validation criteria
After the fixes are implemented, run another audit with the same core settings. Compare:
- Total crawled URLs
- Indexable URL count
- Error URLs
- Redirect chains
- Canonical errors
- Orphan pages
- Broken internal links
- Performance results
- Remaining Critical and High priority Hints
Common Sitebulb Auditing Mistakes
- Using the HTML Crawler when important content depends on JavaScript
- Crawling only internal links and missing orphan pages
- Enabling every report without considering crawl time
- Treating Sitebulb's score as the final SEO diagnosis
- Fixing low-priority metadata before resolving indexability problems
- Assuming every orphan page should be linked or indexed
- Exporting issues without checking affected URLs manually
- Failing to re-crawl after development changes
Final Workflow
Use this sequence for a reliable Sitebulb audit:
- Define the audit scope.
- Create a Sitebulb project.
- Enter the correct start URL.
- Choose the HTML or Chrome Crawler based on how the site renders.
- Enable only the required audit data.
- Add XML sitemaps, Google Analytics, Google Search Console and URL lists.
- Run the crawl.
- Check whether the crawl is complete.
- Review Indexability, Links, On Page and Performance reports.
- Investigate Hints and validate examples manually.
- Prioritise fixes by visibility and business impact.
- Export the findings and re-audit after implementation.
The result is a technical SEO audit with evidence, priorities and clear next actions rather than a list of warnings.