Why Pages Are Crawled but Not Indexed: Mergeflo Guide

Why Pages Are Crawled but Not Indexed: Mergeflo Guide

Short Answer

Short answer: Yes. For a new site, it is normal that some pages are crawled but not indexed while Google evaluates quality and usefulness. This status often reflects duplication, thin or templated content, or weak internal linking. Improve unique value, consolidate duplicates with canonicals, and route internal links to priority pages to earn indexing.

Why New Sites See Crawled but Not Indexed

New domains start with low trust, so Google tests and withholds URLs until pages prove unique value and demand. You will see crawled but not indexed when content is near-duplicate, boilerplate-heavy, light on entities, or orphaned. This is a quality gate, and it is common on sites under 6 months old.

Indexing is a quality gate. If multiple URLs say the same thing, Google will index the strongest one and hold back the rest.

Distinguish status types to pick the right fix. Crawled — currently not indexed means Google fetched the page and paused inclusion. Discovered — currently not indexed means it has not been fetched yet. See definitions and workflows in Google documentation (Google Search Central) and practical breakdowns from Yoast.

Vector infographic showing a left-to-right pipeline for new domains: Discovered to Crawled to an evaluation gate, with some URLs pausing at “Crawled – currently not indexed” and others reaching “Indexed.” Highlights unique value, internal links, and canonicals as signals, using orange, charcoal, and slate brand colors.
Diagram contrasting crawl vs index signals on new domains

What to Change First on a New Site

Tighten quality and routing before you scale URLs or request indexing. Consolidate near-duplicates with rel=canonical, then add 200-300 words of unique, intent-fulfilling content per page: FAQs, specs, examples, or data. Link from already-indexed pages using descriptive anchors, and make sure the page sits in the XML sitemap.

Run a fast, operator-grade triage. In Screaming Frog, group by template and title to spot duplicates. In Ahrefs or SEMrush, confirm target intents and overlapping keywords. A 3-person growth team with a $2k/mo content budget should stabilize 1-2 templates and 20-40 URLs first. Mass-publishing 200 similar pages before validation slows indexing.

Minimal vector site map showing indexed hub pages linking to highlighted priority pages, with dotted canonical lines from duplicate variants to a primary canonical page, using the brand’s orange, charcoal, and slate colors.
Internal linking map highlighting priority pages and canonicals

Status Patterns and Actions on New Sites

Match each Search Console status to a template-level fix and a realistic re-crawl window. Avoid blanket tactics; raise unique value and link equity where it matters.

Index Status Patterns on New Sites: Root Causes, Fixes, and Timelines

Status Pattern Likely Cause On New Sites What To Change Expected Time To Index
Crawled — currently not indexed Near-duplicate or thin templated page Consolidate with canonical, add unique sections 2-8 weeks post-update
Discovered — currently not indexed Weak signals, low priority in crawl queue Add internal links, include in XML sitemap 1-4 weeks after links
Duplicate, submitted URL not selected canonical Multiple URLs target same intent Normalize URLs, set canonicals, remove params 2-6 weeks
Alternate page with proper canonical Correct canonical exists elsewhere Remove duplicates from sitemap, noindex alternates No change expected
Soft 404 Low-value listing or thin page Add substance, real inventory, FAQs, specs 2-6 weeks
Redirect or blocked by robots Technical gating Fix robots, ensure 200 status, resubmit sitemap 1-2 weeks

Reference deeper diagnostics and examples from Marie Haynes for nuanced edge cases.

Expect patterns in Coverage and Page indexing. If more than 60 percent of submitted URLs show Discovered, currently not indexed in the first 30 days, trim your sitemap to the top 200 to 500 URLs, remove redirected or parameter variants, and verify each remaining page returns a fast 200 with TTFB under 800 ms. For Crawled, currently not indexed, add 2 to 5 contextual internal links from already indexed pages, rewrite overlapping copy so at least 30 percent is unique, and ensure only one 200 variant exists with a self referencing canonical. Soft 404 and Thin content cases need richer on-page elements, unique data, and clear intent.

From Fixes to a System with Mergeflo

Turn one-off indexing wins into a repeatable pipeline that scales past 40 URLs. Manual triage works at 20-40 pages; it breaks as you approach 200+ pages because indexing lag compounds and templates drift. You need template-level quality audits, internal link routing, schema, and scheduled refreshes shipped into your CMS.

Mergeflo is an AI search visibility platform for startups. It is an autonomous SEO + AEO content engine: research to published, AI-citable pages in the customer's CMS, with schema, internal links, and ongoing refresh. It is end-to-end and autonomous — it measures AND fixes visibility across Google and AI engines, startup-priced at 149-649 USD per month. If you are dealing with both statuses, pair this system with targeted fixes for discovery and evaluation using our deep dive on how to fix Discovered — currently not indexed in Google Search Console. For evaluation holds, see tactical causes in why a website is crawled but not indexed in Search Console.

Turn fixes into a loop. In Mergeflo, maintain a URL registry with index intent, lastmod, canonical target, content length, and similarity score. Pull URL Inspection API samples and Search Console coverage data daily, then log counts by status. Rules drive work: if Crawled, currently not indexed exceeds 100 URLs for 7 days, open a task to place three contextual links from your 20 most crawled pages and refresh the related sitemap. If a page is under 300 words or over 80 percent similar to another, queue consolidate or noindex. Segment sitemaps under 50,000 URLs or 50 MB, update lastmod on real edits, and watch median time to index.

Frequently Asked Questions

Anchor actions to signals from GSC, sitemaps, and internal links instead of guessing.

How Long Should I Wait Before Escalating a Crawled but Not Indexed URL?

Give updated pages 2-8 weeks after substantive changes. That window assumes you improved unique content, added contextual internal links, and corrected canonicals. If nothing moves by week 8, treat it as a template-level issue and consolidate or deprecate.

Should I Use URL Inspection to Request Indexing for Many Pages?

Use it sparingly for priority URLs after real changes. Mass requests without upgrades rarely move quality-gated pages. Improve content and links first, then request indexing for 5-10 top URLs to validate the template fix.

Does Publishing More Pages Help a New Site Index Faster?

Volume without quality signals slows indexing. Expand only when each template can ship unique sections, entity schema, and internal links from indexed hubs. A small team should stabilize 1-2 templates before scaling to 100+ URLs.

How Do I Tell if This Is a Quality Issue or Crawl Budget?

If Google fetched the page and paused, it is usually quality or duplication. If it remains Discovered — currently not indexed, raise priority with internal links and sitemap placement. Crawl budget limits are rare on small sites unless blocked by robots, noindex, or slow responses.

Turn indexing triage into a workflow. Mergeflo operationalizes keyword-to-cluster publishing with schema, internal links, and ongoing refresh so new sites earn indexing and AI citations fast.

Try Mergeflo →