How to Map Redirects During a Site Migration: Exact vs Semantic Matching
Build a migration redirect map with exact URL matching first, then use a documented Screaming Frog workflow to find semantic matches for changed pages.

A site migration breaks the URLs your visitors and search engines already know.
A redirect map is the spreadsheet that connects every old URL to the best new URL, so link equity and returning visitors land somewhere useful instead of on an error page.
Most teams start with exact matching: the old URL and the new URL point at the same page, just under a different address.
That works until the migration also restructures content.
When pages are merged, split or rewritten, exact matches stop existing, and that is where semantic matching earns its place.
This article walks through both methods, when each one fits, and how to combine them into a redirect map you can actually review before launch.
Start with the exact match pass
Exact matching means you compare the old URL set against the new URL set and pair pages that are genuinely the same content at a new address.
Typical cases include a domain change, a switch from one CMS to another with the same structure, or a trailing slash and protocol cleanup.
The practical steps are simple:
- Crawl the old site to capture every URL worth preserving, along with its title and content.
- Crawl the new site, or the staging environment, the same way.
- Match URLs on shared path segments, slugs or identifiers where they exist.
- Review the pairs manually before you treat them as final.
Automated slug matching gets you a long way, but it is not a review.
Two pages can share a slug and cover different topics, especially after a content restructure.
Treat every automated pair as a proposal that a person confirms.
For pages with no obvious counterpart, resist the temptation to guess.
Leave them unmatched for now and move them into the semantic pass.
Where exact matching runs out
Migrations rarely preserve one-to-one URLs.
Content teams merge thin pages into guides, split long resources into sections, retire product lines and rename categories.
After those changes, an old page may have several plausible new homes, or none at all.
This is the gap semantic matching addresses.
Instead of comparing URL strings, it compares what the pages actually say.
Pages are converted into vector embeddings, which are numeric representations of meaning, and the closest matches are surfaced by similarity.
Screaming Frog's SEO Spider can identify semantically similar pages using LLM embeddings (Screaming Frog).
It is not purpose-built for redirect mapping, but it can help find the closest matching neighbour when redirecting old URLs to new equivalents.
The key phrase is “closest matching neighbour”.
Semantic matching finds candidates.
It does not certify that a redirect is correct.
A page about pricing tiers may be semantically close to a page about plan features, and only a person who knows the content can decide whether that redirect serves the visitor.
How semantic matching works in practice
The following workflow is documented by Screaming Frog for its SEO Spider tool (Screaming Frog); the individual steps below all come from that guide.
- Connect to an AI provider for embeddings, such as OpenAI, Gemini or Ollama, and supply an API key.
- Enable the embeddings prompt from the library and configure the tool to store page HTML so the text is available for embedding.
- Crawl the old and new sites together in list mode, then run crawl analysis to populate the semantic similarity filters.
- Use embedding filter rules so pages from the old site are only matched against the new site, not against each other.
Two configuration details matter for migration work specifically.
If you crawl a staging site blocked from indexing, disable the setting that only checks indexable pages for semantic similarity.
Otherwise, your staging pages will be excluded from matching.
And if you want matches for paginated pages, the relevant configuration can be disabled as well.
Note that this feature requires a paid licence for the software, and the vendor is explicit that it is not the right fit for every scenario.
If your migration is a straightforward domain change with identical URLs, you may never need embeddings at all.
Deciding what to do with weak matches
Semantic similarity produces a ranked list of candidates, not a verdict.
Sort old pages by the strength of their best match and work through them in tiers:
- Clear match: the new page covers the same topic for the same audience. Confirm and add it to the map.
- Partial match: the topic overlaps but the intent differs, for example an old blog post whose content now lives inside a broader guide. Redirect to the guide section if one exists, or to the guide itself.
- No sensible match: the content was retired deliberately. Redirect to the closest relevant category or resource, or let the page go. Redirecting everything to the homepage is a poor experience and a poor signal.
Document the decision for each tier.
When someone asks later why an old URL points where it does, the map should answer the question on its own.
Building the map so it can be reviewed
A redirect map is a working document, not just a CSV you upload at launch.
Give it columns that support review:
- Old URL
- New URL
- Match method (exact, semantic, manual, none)
- Confidence or reviewer notes
- Status (proposed, approved, implemented, verified)
The match-method column is what makes review scalable.
A reviewer can spot-check the exact matches quickly and spend their attention on the semantic and manual rows, where mistakes are more likely.
If a semantic match looks wrong, change the target or downgrade it to a broader destination rather than shipping a redirect that misleads visitors.
Keep the old-site crawl data alongside the map.
If a dispute comes up about what a page used to cover, you can check the stored content instead of relying on memory.
Comparing crawls over time is also how you verify redirects after launch.
The same crawl-and-compare discipline applies to post-migration checks, as described in how to compare website crawls over time.
Common failure modes to avoid
- Matching on slugs alone. Identical slugs across sites can hide different content. Confirm the page, not just the address.
- Treating similarity scores as truth. A high similarity means the text is close, not that the redirect is right for the visitor.
- Redirecting everything to the homepage. When no good target exists, choose the nearest relevant section, or accept that some URLs will not be redirected.
- Skipping the staging configuration. If your semantic tooling only checks indexable pages, a non-indexable staging site will silently produce no matches.
- No post-launch verification. A map that was never tested against the live server is a plan, not a result. Crawl the old URLs after launch and confirm each one resolves to its intended target.
Choosing between the two methods
The choice is not either-or.
Exact matching is faster, cheaper and more reliable wherever the content genuinely survived the migration unchanged.
Semantic matching fills the residue: restructured, merged and rewritten pages where URL patterns tell you nothing.
A sensible sequence for most migrations:
- Crawl both sites and build the exact match set from shared paths and slugs.
- Review those pairs; they are usually safe but not automatic.
- Run semantic matching on the unmatched remainder, using the old-to-new filter rules described above.
- Review semantic candidates by tier and record decisions.
- Handle true orphans with a deliberate destination or a documented decision to drop them.
- Verify every mapped redirect after launch with a fresh crawl.
Matching records by meaning rather than by string also appears in revenue operations, such as deduplicating contacts across systems.
The tradeoffs are similar, as covered in record matching.
What good looks like at launch
Before cutover, every row in the map should have an approved status and a named reviewer.
After cutover, crawl the old URL set and confirm three things.
No old URL should return an error, each redirect should land on its mapped target, and no redirect should chain through extra hops.
Any surprises go back into the map as new rows, with the fix recorded.
A redirect map built this way does more than preserve rankings.
It preserves the reasoning behind every redirect, which is what lets the next person maintain the site without starting over.
Match on content, not just on URLs, and record why every redirect points where it does.
How Meshline can help. Connect automation, Organic Marketing (demand generation), and customer lifecycle management (Revenue Intelligence).
Bring topic planning, content publishing and performance feedback into the conversation about your workflow. Book a Meshline demo.