Skip to content
TrustList
Blog

How we cleaned up spam, thin pages and dead links on TrustList in 2026

Editorial

By TrustList Editorial

How we found and removed spam, thin profiles, paid guest posts and dead website links in September 2026, why we archive instead of delete, and what is unfinished.

About How we cleaned up spam, thin pages and dead links on TrustList in 2026

How we cleaned up spam, thin pages and dead links on TrustList in 2026

TrustList came back after about a year offline as an ordinary directory: a large catalogue built by import, with plenty of pages nobody had read since. On 25 September 2026 our AdSense application was refused with the words "Low-value content". Around 17 September, 24,350 of our live listings, about a quarter, carried fewer than 50 words of profile text, and 82 per cent of the older company profiles with 50 words or more (54,541 of 66,473) were written in the business's own voice.

Within days we set out to turn that into a curated catalogue. Between 26 and 27 September we archived 164 wholly spam listings, 24 partly spam ones and 52 paid guest posts, rewrote 6,165 thin listings from the business's own site, and took 2,675 dead or parked domains offline. By the afternoon of 29 September, 29,722 of 72,564 live company listings had been rewritten to one written standard.

This article is the detailed account of that clean-up: what we found, what we removed, what we only switched off, and what we are still working through. Some of it is not flattering, and we would rather say so. Every figure carries the date it was true, so please read the numbers as snapshots. The story of how our search traffic behaved over the same period is told in our earlier piece on search traffic, and we will not repeat it here.

What the reviewer saw and what we counted

On 25 September 2026 we counted what we were asking Google to judge. Our main sitemap offered 99,592 listing pages, 18,333 ranking pages, 1,000 internal search pages and about 960 other pages: articles, topic rankings, checklists and the like. The listing pages were the problem. A quarter of them said almost nothing, and most of the rest were the business's own website copy, imported years earlier.

The numbers from that week: about 24,350 live listings carried fewer than fifty words of profile text, and some had none at all. Of roughly 74,000 older listings with fifty words or more, 82 per cent of the company profiles (54,541 of 66,473) were written in the business's own voice ("we", "our"), and 9,413 still carried website furniture such as "read more" and "cookie". Only around 1,130 listings had been written or rebuilt by us. We had 155 editorial articles and 61 guest posts.

The rankings had a parallel weakness. Many were assembled from the same thin listings, and they earned very little: 255 thousand impressions and 67 clicks over the period we measured.

Google's guidance on creating helpful content and the AdSense programme policies ask a simple question of a site like ours: is there original value here that a visitor could not get by reading the business's own website? For a large share of the catalogue, honestly, there was not.

We set ourselves a rule, in the spirit of our listing guidelines. Every thin listing ends in one of three states. Either we rebuild it in our own words from what the business itself publishes, or we retire it because the business or product has ended, or we take it offline because nothing can be found to write from.

Spam that had taken over old domains

The most uncomfortable finding came from a full-text scan of every live listing, 89,840 of them, on 26 September 2026. We searched for the signals of spam: adult content, escort and dating text, casino and slot spam in several languages, pills, payday loans, love spells, essay mills, "crypto recovery" pitches, parked and for-sale domains, placeholder pages, legal-page text and non-Latin scripts.

The scan flagged 1,719 listings. A keyword hit is not a verdict, so six human reviewers read every flagged listing in context against a written brief. Their classification:

  • Spam, 164 listings wholly and 24 partly. The text belonged to a domain taken over after the business let it lapse. A recruitment agency whose listing text was adult listings. An escort advertisement on a marketing agency. Gambling link lists on venture capital firms. A love-spell advertisement with a phone number and threats.
  • Illicit, 12 listings. The businesses themselves: essay mills and assignment writers, a crypto "recovery" service and a fortune-teller. Contract cheating is an offence in England, and we do not want to send readers to it. Nine had claimed their own listings.
  • Junk, 432 listings. No description at all: parked domains with asking prices, "this site can't be reached" text, lorem ipsum, unedited theme demos, and policies stored where a description should be.
  • Fine, 1,087 listings. The keyword was a false positive. Vendors of software for the gambling industry, real firms writing in their own language, a "coming soon" note about one feature.

That last number matters. Nearly two thirds of what the scan flagged was legitimate. Had we removed on the keyword alone, we would have wrongly deleted more than a thousand honest businesses. Reading each one was slow, and it was the only fair way.

On 26 September 2026 we removed the 176 wholly spam or illicit listings (164 companies and 12 products). The 24 partly-spam listings were treated differently. These were real businesses with injected text, so we held them back to be rebuilt rather than removed.

Paid guest posts and reviews that should never have shown

We also had to look at our own blog. Of the 61 guest posts, 51 were promotional: backlinks to the author's own firm, "hire us" pages and press releases. One was a spun article. Separately, one article published as editorial was an unfinished generator stub. It told the reader to set an API key "for a fully written article" and then listed raw scraped blurbs. That one was ours, and it was embarrassing.

Google's spam policies describe hosting third-party promotional pages to borrow a site's reputation as site reputation abuse. Our guest posts were exactly that. They had carried a noindex tag since 25 September, but they were still live and linked. On 26 September 2026 we archived 52 of them, the 51 promotional ones and the spun one, along with the stub.

Reviews got the same look. Of 17 published reviews, seven were hidden: a "funds recovery" scam pitch, a gambling-software review on an app developer, salad-shop copy on a recruitment agency, an essay-mill testimonial posted from the mill's own domain, an empty review, and two businesses' own marketing copy posted as reviews of themselves.

Then we found a bug: blocked reviews were never actually hidden. The filter used an operator our content system does not understand, so it filtered nothing, and a blocked review still appeared with its markup. That was fixed the same day, and blocking a review now recalculates the listing's ratings and Trust Score.

Thin listings: written from the source, or parked

Fast, plain rewriting was the biggest single piece of work. On 26 September 2026 the queue held 25,018 live thin listings. The owner's instruction was to give them all a fast pass: read each business's own website, write a short third-person profile from what it says, and put it live.

Here is how the pipeline worked, in reader terms.

A polite reader visited each listed website once and classed it as working, dead, empty, moved or parked. Sites that draw their content with scripts were opened in a real browser, and sites behind bot walls were left alone; we never bypass them. AI-assisted drafting then wrote each profile under strict rules: third person, 110 to 180 words, facts only from the page, no hype, no prices, no contact details in the text. An automated check refused any draft with the wrong length, first-person voice, hype words, prices, numbers or names not in the source, long runs copied from the page, or a missing business name in the opening. Batches went to a copy of the database first, then live, each with a record that could reverse it.

By the end of that programme, 6,165 listings had been rewritten and were live. For 4,739 of them we also stored the email, phone or social links the business publishes on its own site, but only where our record was empty. Every rewrite is marked with the date it was rebuilt, and that mark is what makes a listing count as written by us.

In 314 cases a business had moved to a new domain and the writer judged it the same business, so the listing's website followed it. A further 104 rebrands were wrongly parked because a first draft failed the checker; we restored them.

The other 15,238 thin listings could not be written from anything, and we parked them. Parking means archiving: the page returns a 404 and disappears from rankings and sitemaps, but nothing is deleted. The reasons were:

The biggest groups were 9,392 with no website on the listing and 3,112 whose website did not answer. The rest were sites with nothing readable (809), sites that were not this business (496), domains now owned by someone else (285), parked or for-sale domains (382), sites with too little to describe the business (399), sites about something else (170), out-of-scope listings (97), bot walls (63) and 33 whose site said the business had closed.

We did not park IT service providers (held for a separate experiment until 13 October 2026), claimed listings, or any listing with a review.

Why we archive and do not delete

The choice of archiving over deleting was deliberate, and we made it several times over.

First, we make mistakes. The 104 rebrands are a live example: parking is a flag with a record of exactly which rows were touched, so we undid it in one step. A delete would have needed a restore from backup.

Second, a business might come back. A listing with no website today may get a verified one next month, and some of the 9,392 lost a wrong website in an earlier repair.

Third, a business that says it has closed, or a product that has ended, gets a retirement notice instead. It stays readable but is kept out of search and out of every ranking. Deleting would lose the fact that the business existed and stopped. The rule we keep coming back to is that every removal must be reversible.

Dead and parked websites: flagged, not linked

The other 77,998 live listings that were not thin had a different problem: their outbound links might be dead. On 27 September 2026 we checked every listing's website.

On 27 September 2026, of the sites not served from a large content network, 57,228 worked, 2,702 had no DNS record, 629 were parked, 3,874 were dead, 4,655 had moved, 992 were alive behind a bot wall and 2,300 were undecided after two timeouts.

Sites served from that content network were left alone in this pass, since their errors are often temporary.

The owner's decision was to remove what should be removed. Two things followed.

  • 2,675 listings went offline, the ones whose domain had no DNS record or was parked. They were not on the content network, were not IT service providers, were not claimed, and had not been rebuilt by us.
  • 4,530 listings were flagged instead, covering 5,008 contact records: 4,350 dead, 574 with no DNS record and 84 parked. Some flagged listings were kept online deliberately: claimed listings, ones we had rebuilt, and IT service providers, whose turn comes after 13 October 2026.

A flagged listing keeps its profile, which is our writing, but no longer links to a website that cannot be reached. The contact card shows the address struck through, with the date we checked it.

We will re-check the dead ones in a week, and any still dead go offline. The 6,575 moved sites are next: the owner wants a record of the previous name kept, and where a different business took the site over, the listing goes offline.

Ranking pages that exist only when a ranking backs them

Our ranking pages had a structural problem that had nothing to do with the quality of text. The system made a "best" page for every combination of category and term that had at least one listing. Search Console's coverage export of 28 September 2026 showed what that produced: 90,416 pages indexed, against 121,433 crawled but not indexed, 34,867 discovered but not indexed and 19,689 soft 404s. Indexed pages had fallen from 139,511 on 25 August. Daily impressions had dropped from 22,638 to about 4,000 or 5,000.

In our production data, 68,678 category-and-term pairs had at least one live listing, but only 11,962 had five or more listings that were not thin. Some pages were plainly silly, such as accounting firms "for WordPress".

The owner's instruction was direct: no data, no page. So a page in the "best" section now exists only when all three of these are true:

  1. Every term in the address is one we allow people to browse.
  2. A published ranking entry names that exact combination.
  3. At least one live listing of that type carries the term.

Otherwise the page answers 404, and nothing on the site links to it. A term chip on a listing page becomes a plain label unless its ranking exists.

We seeded the allowed list, on 28 September 2026, from the pairs that already passed the five-real-listings test: 11,902 pairs plus 20 pages from our search-ranking experiment. The list grows by itself overnight as new pairs reach the bar. Taking a page down is a matter of unpublishing its entry, and the nightly job will not recreate it.

If a refresh of the list fails, the last good copy stays in use, so a database hiccup cannot turn the rankings into errors. This does not solve everything: pages such as accounting firms for WordPress still cleared the bar, so junk pairs that pass the count remain a separate step.

Crawl waste: links that only lead to redirects

The last piece concerned how our pages talk to crawlers. The coverage export showed 34,786 pages listed as alternates of a proper canonical page. Nearly all were sorted and paged variants of ranking pages.

The cause was simple. Every link on a ranked page carried both a sort choice and a page number, defaults included, so one ranking offered Google around 70 addresses that all pointed to one canonical page.

We changed this on 28 September 2026. Default values are left out of the address, so page one is the bare page. Sort choices are marked so crawlers do not follow them, and so are paging links inside a sorted view. Pages two onward in the default order stay ordinary links.

On listing pages, "Write a review" and "Own this listing? Claim it" led crawlers only to sign-in redirects or blocked pages. The export showed 28,908 "page with redirect" entries and 10,041 blocked by robots. Those links are now nofollow too, and a reader sees no difference. Google keeps the old addresses until it recrawls them, so we expect months rather than days.

Before and after, in numbers

Here is where things stood before the clean-up and where they stood once it had run. Each pair carries its date, and each is a snapshot rather than a final figure.

  • Thin profiles: around 17 September 2026, 24,350 live listings had fewer than 50 words of profile text -> on 26 September, 6,165 of them had been rewritten from the business's own site, and 15,238 with nothing to source were parked (archived, reversible).
  • Voice: around 17 September, 54,541 of 66,473 older company profiles with 50 words or more (82 per cent) spoke as "we" and "our" -> on 29 September, 29,722 of 72,564 live company listings had been rewritten to the standard, with about 19,000 more queued to publish that day.
  • Listings in our own words: on 28 September, about 11,500 of 81,127 live, active listings (14 per cent) -> the 29 September figure above.
  • Spam and paid promotion: before 26 September, 164 wholly spam listings, 24 partly spam ones and 52 paid guest posts live -> all archived on 26 September, and 7 reviews blocked.
  • Dead links: before 27 September, listings pointed to domains with no DNS record or parked domains -> 2,675 listings taken offline, and 4,530 flagged so their pages never link a dead site.
  • Ranking pages: on 25 September, 18,333 ranking pages in the sitemap, many built from import tags, and 68,678 category-and-term pairs with a page but only 11,962 with five or more listings that were not thin -> on 28 September, a /best page exists only when a published ranking backs it, about 12,000 are published, and category pages without one answer 404 with nothing linking to them.

Every rewritten profile now follows one standard: what the business offers, who it works with, and key facts. It uses only facts from the business's own site and our records, and it passes automated checks for length, copying, hype, invented numbers or names, prices and contact details.

For a reader, the change is that a page on TrustList is now less likely to be an import nobody looked at. A profile is more likely to be written by us in the third person, a website link is less likely to lead nowhere, and a ranking page exists because a ranking stands behind it. We are not claiming the catalogue is finished, and the figures above show how much was thin to begin with. What we can say is what has changed, on which day, and by how many listings.

Sources

Figures about TrustList come from our own internal records, each dated in the text above (17 to 29 September 2026). They describe the state on those days and will have moved since.