Internal Linking: Automation That Works
Imagine: a site with 300 articles, every second one is an orphan (has no internal links). Google can't find these pages, PageRank doesn't flow, users get lost. Manual linking at that scale takes 3 editors a week — and a month later 10 new articles appear, so the process repeats. Our experience shows: an automated system pays for itself in 1–2 months.
We take on the design and implementation of such a system — turnkey, tailored to your stack. Below is how we do it.
Automation Solves the Orphan Page Problem
Orphan Pages Lose Traffic and Weight
Without incoming internal links, crawlers rarely find them, and if they do, they don't pass link equity. The result: low rankings, small reach. On one project we found 40% of blog posts had no internal links — after automation, traffic to them grew 2.5× in 3 months.
Manual Linking Doesn't Scale
At 500+ pages, an editor can't remember all relevant connections. They insert links intuitively, missing obvious pairs. An automated system (keyword + semantic) finds links a human would overlook, in seconds.
Anchors Lose Relevance
Manually, it's hard to ensure the anchor exactly matches the target page's content. Keyword-matching guarantees: the link is placed on the exact word that describes the page. Semantic matching goes further — it finds pages by meaning even if the anchor doesn't match.
Why a Hybrid Approach Outperforms a Single Method?
We use a hybrid: keyword-based for direct matches, semantic matching for the "Related Articles" block. This combination yields 30–60% more clicks compared to using just one method.
Keyword-Based Linking
We build a keyword dictionary from titles and meta fields of all published articles. We sort by length so long keys match before short ones (avoiding partial matches like "React" inside "React Native").
class AutoLinker
{
private array $linkMap;
public function __construct()
{
$this->linkMap = Cache::remember('autolink_map', 3600, function () {
return Article::where('is_published', true)
->get()
->flatMap(fn($a) => collect($a->keywords)->mapWithKeys(
fn($kw) => [$kw => route('articles.show', $a->slug)]
))
->all();
});
uksort($this->linkMap, fn($a, $b) => strlen($b) - strlen($a));
}
public function process(string $html, string $currentUrl): string
{
$dom = new \DOMDocument();
@$dom->loadHTML(mb_convert_encoding($html, 'HTML-ENTITIES', 'UTF-8'));
$linked = [];
foreach ($this->linkMap as $keyword => $url) {
if ($url === $currentUrl) continue;
if (isset($linked[$url])) continue;
$xpath = new \DOMXPath($dom);
$textNodes = $xpath->query('//text()[not(ancestor::a) and not(ancestor::code) and not(ancestor::pre)]');
foreach ($textNodes as $node) {
$pattern = '/\b' . preg_quote($keyword, '/') . '\b/ui';
if (preg_match($pattern, $node->nodeValue)) {
$new = preg_replace($pattern,
"<a href=\"{$url}\">{$keyword}</a>",
$node->nodeValue, 1
);
$fragment = $dom->createDocumentFragment();
@$fragment->appendXML($new);
$node->parentNode->replaceChild($fragment, $node);
$linked[$url] = true;
break;
}
}
}
return $dom->saveHTML();
}
}
After processing, each page gets 3–5 additional internal links that distribute weight evenly.
Semantic Linking with Vector Embeddings
For the "Related Articles" block, we use vector embeddings. When an article is saved, we generate an embedding via OpenAI text-embedding-3-small and store it in PostgreSQL with the pgvector extension. Finding similar articles is a cosine distance query.
class SemanticLinker
{
public function findRelated(Article $article, int $limit = 5): Collection
{
return Article::selectRaw('*, embedding <=> ? AS distance', [$article->embedding])
->where('id', '!=', $article->id)
->where('is_published', true)
->whereRaw('embedding IS NOT NULL')
->orderBy('distance')
->limit($limit)
->get();
}
public function generateEmbedding(Article $article): void
{
$text = $article->title . "\n" . strip_tags($article->excerpt);
$response = Http::withToken(config('openai.key'))
->post('https://api.openai.com/v1/embeddings', [
'model' => 'text-embedding-3-small',
'input' => $text,
]);
$embedding = $response->json('data.0.embedding');
$article->update(['embedding' => json_encode($embedding)]);
}
}
Result: the related articles block isn't based on tags alone but on real semantic similarity. Click-through rate on these links is 30–60% higher.
Related Articles React Component
A ready-made UI component that loads data from an API and displays as cards. Requires React 18+ and TanStack Query.
// RelatedArticles.tsx
interface Article {
id: number;
title: string;
slug: string;
excerpt: string;
category: string;
}
export function RelatedArticles({ articleId }: { articleId: number }) {
const { data: related } = useQuery({
queryKey: ['related', articleId],
queryFn: () => fetch(`/api/articles/${articleId}/related`).then(r => r.json()),
staleTime: 5 * 60 * 1000,
});
if (!related?.length) return null;
return (
<aside className="mt-12 border-t pt-8">
<h3 className="text-lg font-semibold mb-4">Related</h3>
<div className="grid grid-cols-1 sm:grid-cols-2 gap-4">
{related.map((article: Article) => (
<a key={article.id} href={`/articles/${article.slug}`}
className="block p-4 border rounded-lg hover:border-blue-400 transition-colors">
<span className="text-xs text-blue-600 uppercase tracking-wide">{article.category}</span>
<h4 className="font-medium mt-1 text-sm leading-snug">{article.title}</h4>
</a>
))}
</div>
</aside>
);
}
What Results Does the Hybrid Approach Show?
Compare key metrics before and after implementation:
| Metric | Before | After |
|---|---|---|
| Internal links per page | 0–2 | 5–8 |
| Orphan page share | 40% | < 5% |
| CTR on related articles block | 8% | 14–18% |
| Page load time (LCP) | 2.3 s | 2.5 s (negligible) |
Case study: an e‑commerce site with 500 products
After implementing hybrid linking, category traffic grew 35% in 2 months, and average pages per session increased from 1.8 to 3.2. The system paid for itself in 1.5 months.What’s Included
| Component | Description | Duration |
|---|---|---|
| Keyword-based linker | Dictionary collection, HTML replacement, duplicate prevention | 1–2 days |
| Semantic linker | Embedding integration, pgvector, API | 3–4 days |
| Related articles component | React/Vue component, API routes, caching | 1–2 days |
| Linking report | SQL queries to monitor orphans, top incoming links | 0.5 day |
| Documentation and training | Process description, admin access | 0.5 day |
Total: 3–7 business days depending on your stack and integration complexity. Contact us for a preliminary estimate of your project.
How We Work
- Analysis — audit current structure, collect semantics, identify orphan pages.
- Design — choose approach (keyword, semantic, or hybrid), dictionary architecture, CMS integration.
- Implementation — write linker code, set up embeddings, develop UI component.
- Testing — verify on test data: no layout breakage, no circular links, correct exceptions.
- Deploy and monitor — roll out to production, log errors, track metrics (link count, LCP, INP).
Checklist: Are You Ready for Automated Linking?
- [ ] All articles have unique titles and meta descriptions
- [ ] CMS has a "keywords" field (or we can add it)
- [ ] Your stack supports PHP 8.3+ or Node.js (for embeddings)
- [ ] You have PostgreSQL (for pgvector) or are ready to switch
If all items are checked — implementation takes 3–5 days. If not, we’ll help refine the structure.
Book a free consultation — we will analyze your site, provide an estimate, and give precise timelines.







