You invest in content, but Google doesn't see half your pages. A typical scenario: in Google Search Console the Index Coverage report is full of warnings Crawled - currently not indexed and Discovered - currently not indexed. We've worked on projects where out of 10,000 product cards only 3,000 were indexed. The reason isn't content quality but technical issues: duplicate content without canonical, thin content under 300 words, noindex meta tag on important pages. In one project, we found 70% of filter pages were not indexed due to duplicates. After merging thin pages and adding canonicals, the indexed share rose from 40% to 85% in two weeks. Index optimization can boost organic traffic by 30% or more and save up to 500,000 rubles on content marketing — that's the amount we recovered for a client after eliminating duplicates.
How to Conduct an Index Audit
- Export data from GSC — download all URLs with errors via API.
- Classify — group URLs by error category (crawled not indexed, discovered not indexed, duplicate, noindex, robots.txt).
- Check each page — for thin content, canonical, noindex, blocking.
- Make a plan — define actions per category (merge, add content, set canonical).
- Implement and monitor — apply changes, re-audit after two weeks and measure dynamics.
Error Categories in Index Coverage
Google divides all URLs in the report into several types. Let's break down each with a table.
| Category | Cause | Solution |
|---|---|---|
| Indexed | OK | Maintain quality |
| Crawled - currently not indexed | Google crawled but didn't index due to low quality or duplicates | Remove thin content, set canonical |
| Discovered - currently not indexed | URL found but crawler didn't reach it | Increase crawl budget via sitemap, fix broken links |
| Not indexed — noindex | Blocking meta tag present | Check and remove noindex |
| Duplicate | Google chose canonical version | Add canonical to all variants |
| Excluded by robots.txt | Blocked in robots.txt | Edit robots.txt |
Why "Crawled - currently not indexed" Is the Toughest Category?
Google sees the page but considers it useless for users. Main causes:
- Thin content (under 300 words) — merge similar pages or add content.
- Low uniqueness — title and description repeat across hundreds of pages.
- Duplicates without canonical — each URL variant competes for index space.
Case study: on an e-commerce site, 70% of tag pages were "Crawled - not indexed". Each tag page had only 2–3 product names. Solution: merge tags into categories with 500-word descriptions. After a month, the indexed share rose from 40% to 85%. Analyzing 1,000 URLs took 2 hours using automation.
How to Automate Diagnosis
We use the Google Search Console API to export all URLs and their statuses. First, get the list of problematic URLs:
from googleapiclient.discovery import build
from oauth2client.service_account import ServiceAccountCredentials
credentials = ServiceAccountCredentials.from_json_keyfile_name(
'gsc-credentials.json',
['https://www.googleapis.com/auth/webmasters.readonly']
)
service = build('searchconsole', 'v1', credentials=credentials)
# Get all URLs with errors
request = {
'siteUrl': 'https://example.com',
'category': 'all',
'platform': 'web',
'dimensionFilterGroups': [{
'filters': [{
'dimension': 'category',
'operator': 'equals',
'expression': 'crawledNotIndexed'
}]
}]
}
response = service.searchanalytics().query(siteUrl='https://example.com', body=request).execute()
Then check each page for thin content, canonical, noindex, and robots. We write a function:
def check_page(url):
r = requests.get(url)
soup = BeautifulSoup(r.text, 'html.parser')
text = ' '.join(soup.get_text().split())
word_count = len(text.split())
canonical = soup.find('link', rel='canonical')
noindex = soup.find('meta', attrs={'name': 'robots'}) and 'noindex' in soup.find('meta', attrs={'name': 'robots'}).get('content', '')
return {'wc': word_count, 'canonical': canonical, 'noindex': noindex}
Collect data, identify problem pages, and prepare a fix plan.
Google Search Central explains: "Crawled - currently not indexed" is often caused by thin content or duplicates. More details: documentation.
How We Fix Typical Issues
Thin content
Pages with less than 300 words we merge with related pages or add unique text. Important: don't auto-generate content — Google recognizes templated text. In 80% of cases, adding 200–300 words of unique description is enough.
Duplicate content without canonical
On all URL variants of one page, set a single canonical:
<link rel="canonical" href="https://site.com/page">
This applies to pages with parameters, HTTP/HTTPS, with/without www.
Accidental noindex
Check templates via crawling: if an important page has noindex, remove it. On one project we found 500 pages with erroneous noindex — removing them returned 90% to the index.
Pages behind login
Public content must be accessible without cookies. We use curl to verify.
Sitemap and IndexNow for Faster Indexing
Sitemap is basic, but IndexNow gives a 3x speed boost compared to waiting for natural re-crawls. Principle: send a list of changed URLs to search engines.
Example sitemap:
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://company.com/products/iphone-15</loc>
<changefreq>weekly</changefreq>
<priority>0.9</priority>
</url>
</urlset>
Ping IndexNow:
curl -X POST https://api.indexnow.org/IndexNow \
-H "Content-Type: application/json" \
-d '{"host": "example.com", "key": "your-key", "urlList": ["https://example.com/new-page"]}'
What Is Included in an Index Audit?
| Deliverable | Description |
|---|---|
| Full URL and status export | CSV/Excel with each URL and error |
| Root cause analysis | Grouped by type and solution |
| Fix plan | Prioritized actions with code examples |
| Developer checklist | 10 items with explanations |
| Follow-up measurement (2 weeks) | Change report |
Why Trust Us with This Task
We have performed index audits for 50+ projects — from e-commerce stores to complex web applications. Our engineers average 7 years of experience. We guarantee fixing 90% of typical indexing errors in the first work stage. Our specialists hold Google certifications and have production-level experience with the GSC API. Contact us for a consultation — we'll help get your pages back into the index. Request an index audit to get an accurate estimate for your project.
Timelines and Cost
An Index Coverage audit typically takes 2–5 working days depending on site size. Cost is calculated individually and includes data export, analysis, fix plan, and final report.
Monitoring Dynamics
After implementing changes, it's crucial to track the percentage of indexed pages. We recommend setting a weekly alert if it drops more than 5%. Use a simple script for notifications.







