I run a small blog and I write long, researched posts that take weeks. Last week I searched for one of mine on google.com/ieand found a copy of it hosted on translate.kagi.com sitting above my own page. it wasn't a translation, it was the same English text I wrote, on a kagi.com URL. A few days later my page was gone from the results for that term altogether, and the copy was still there.
Mine: https://returnzero.win/2026/09/01/how-i-came-to-choose-the-feiyue-fye355/
The copy: https://translate.kagi.com/returnzero.win/2026/09/01/how-i-came-to-choose-the-feiyue-fye355
I spent a while assuming the fault was mine. I checked robots.txt, meta robots, response headers, canonical tags, sitemaps, whether Googlebot was being blocked. All clean. My page carries a correct self-referencing absolute canonical. It is indexed. It just loses to the copy with a higher domain authority.

So I went and looked at the copy instead, and here is what I found:
- https://translate.kagi.com/robots.txt is
User-agent: * followed by Disallow: with an empty value. That is the explicit "crawl everything" directive. Nothing restricts the proxy paths.
- The proxied URL returns 200 with no
X-Robots-Tag header.
- There is no rel=canonical pointing back to me. Not in the response, and not injected by the JS bundle. I grepped for it and found only an unrelated dictionary UI string.
- The URLs are path based, translate.kagi.com/<domain>/<path>, so a crawler reads them as ordinary pages on a strong domain rather than proxy output.
That combination means there is nothing telling Google the two pages are related. They compete as separate documents and the one on the stronger domain wins. My canonical is only ever visible on my side of it.
I want to be fair about what is happening here. Nobody scraped me. Someone pasted my URL into Translate, your backend fetched my page once, identified itself honestly as KagiTranslatePreview/1.0, and respected my robots.txt, which is more than most crawlers manage. I dont think this is intentional bad behaviour. It is probably a default that nobody thought about, but the cost of it falls on people like me who have no authority to spare.
The part that should worry you more than it worries me
An empty Disallow: on a path-based proxy means anyone can append any URL and have it crawled and indexed on kagi.com's authority. Google Translate shipped this same defect twice. In 2010 translated pages showed up in Google's own results and the Translate team fixed it once someone flagged it. In December 2023 it came back through translate.google.com/website, which robots.txt was not disallowing, and roughly 4.9 million pages ended up indexed and actively used as a spam vector.
Google could absorb that because Google owns the index. Kagi does not. A subdomain hosting unlimited third-party content with crawling explicitly invited is the structural pattern Google's site reputation abuse policy targets, and enforcement there is a manual action that can apply domain-wide. You are obviously not doing this on purpose, but from the outside the shape is identical, and it only takes one person working out that translate.kagi.com is free indexing on a clean domain.
What I am asking for, in order of preference
- Send
X-Robots-Tag: noindex on proxied page responses. Translated pages do not need to be in anyone's index.
- Pass through the source page's rel=canonical as an absolute URL.
- At minimum, add a Disallow rule for the website proxy paths in robots.txt.
I like what you do, which is part of why this one stings. Small Web exists because you think independent sites are worth finding. Right now a different part of the product is quietly outranking one of them with its copied text.
I only caught this because I went searching for my own post and noticed the URL was not mine. Most people likely never run that search. Every site that has been through Translate since launch is potentially in the same position, and the ones least likely to notice are exactly the small independent sites with no analytics habit and no Search Console account. Fixing the default fixes all of them at once.
References:
https://www.billhartzer.com/google/google-translate-pages-are-crawled-and-indexed-in-google-allowing-spam/
https://searchengineland.com/a-lesson-from-the-indexing-of-google-translate-blocking-search-results-from-search-results-62529
The fix:
Proxied pages at translate.kagi.com should not be indexable by default.
In order of preference:
Send X-Robots-Tag: noindex on proxied page responses. Translated output has no reason to be in a third-party index, and this is the only option that fully removes the competition with source pages.
If the pages need to stay indexable, pass through the source page's rel=canonical as an absolute URL so Google consolidates the copy into the original instead of ranking them against each other.
At minimum, add a Disallow rule for the website proxy paths in robots.txt. This is weaker, because a disallowed URL can still be indexed if linked from elsewhere, but it closes off the bulk crawling case.
Ideally an existing indexed copy would also be retired, since a noindex header only takes effect once Google recrawls and rerenders the page, which can take weeks.