मैंने लगभग 3,000 नए पेजों के साथ sitemap अपडेट deploy किया, जिनकी structure /en/blog/... थी। थोड़ी देर बाद Yandex Webmaster में एक unexpected pattern दिखा: Yandex समान paths को /blog/... के नीचे crawl करने की कोशिश कर रहा था, यानी /en prefix के बिना।
अगर वे URLs सीधे 404 लौटाते, तो हजारों बेकार crawl attempts non-existent paths पर जा सकते थे। अच्छी बात यह थी कि मैंने पहले से language prefix के बिना routes से सही English URLs तक permanent 308 redirects लगाए हुए थे।
यह छोटी defensive routing layer मेरी अपेक्षा से कहीं ज्यादा उपयोगी निकली।
मैंने वास्तव में क्या देखा
- मैंने लगभग 3,000 नए
/en/blog/...पेजों वाला sitemap प्रकाशित किया। - उसके बाद Yandex Webmaster ने दिखाया कि Yandex corresponding
/blog/...URLs crawl करने की कोशिश कर रहा है। - वे alternate paths पहले से 308 redirects से covered थे।
- इसलिए requests 404 पर खत्म होने के बजाय intended
/en/blog/...URLs पर पहुँचीं।
मेरे original description में एक जरूरी correction है: मैं यह साबित नहीं कर सकता कि Yandex ने “sitemap गलत parse किया।” Sitemap अपडेट के बाद मैंने unexpected paths देखे, लेकिन timing अकेले internal cause साबित नहीं करती। Search engines URLs को कई signals और historical sources से discover कर सकते हैं। मजबूत evidence के बिना accurate statement बस इतना है कि Yandex ने ऐसे paths crawl किए जिनकी मुझे उम्मीद नहीं थी।
यह distinction जरूरी है। Crawler का strange URL request करना observation है; उसने वह URL क्यों चुना, यह अलग claim है।
308 redirects क्यों काम आए
मेरी redirect logic shorter path को localized URL का permanent alias मानती थी:
/blog/example-post -> 308 -> /en/blog/example-postइसलिए crawler unexpected URL से आने पर भी उस page तक पहुँचता था जिसे मैं वास्तव में serve करना चाहता था।
308 Permanent Redirect एक permanent HTTP redirect है जो request method और body को preserve करता है। Normal crawler GET requests में method preservation आम तौर पर मुख्य बात नहीं होती; मेरे case में महत्वपूर्ण property यह थी कि redirect explicitly permanent है। Yandex Webmaster की current documentation 301 और 308 दोनों को permanent redirects के रूप में classify करती है।
redirects पर Yandex Webmaster की official documentation.
इसका मतलब यह नहीं है कि SEO के लिए 308 हमेशा 301 से बेहतर है। मेरे पास 308 पहले से थे और उन्होंने वही काम किया जिसकी मुझे जरूरत थी: unexpected URL dead end नहीं बना।
Redirect safety net है, sitemap fix नहीं
Redirects ने impact को contain किया, लेकिन extra crawling को desirable नहीं बनाया। हर unnecessary redirect एक अतिरिक्त request और hop है। बहुत broad redirect rule URL generation की गलती को भी छिपा सकती है, अगर आप bad URLs के source को investigate करना बंद कर दें।
अगर sitemap में खुद outdated या redirecting URLs हैं, तो सही fix sitemap को सुधारना और final URLs देना है। Redirects को old, alternate या accidentally discovered paths को protect करना चाहिए; sloppy URL data को justify नहीं करना चाहिए।
मेरे case में sitemap पहले से /en/blog/... इस्तेमाल कर रहा था। Redirects ने बस site को ज्यादा tolerant बनाया जब crawler /blog/... से आया।
आज ऐसी स्थिति में मैं क्या verify करूंगा
- Actually deployed sitemap खोलूंगा। सिर्फ generator code पर भरोसा नहीं; final file और कुछ real URLs check करूंगा।
- Final responses check करूंगा। Index होने वाले sitemap URLs ideally बिना avoidable redirect chains के सीधे सही page पर जाने चाहिए।
- Predictable alternate paths test करूंगा। अगर old या prefix-less URL का permanent destination है, mapping explicit और one-to-one होनी चाहिए।
- Redirect chains avoid करूंगा।
A -> B -> C,A -> Cसे ज्यादा difficult है। - जहाँ संभव हो crawler reports को server-side evidence से compare करूंगा। Webmaster tools useful हैं, लेकिन हमेशा URL discovery का source नहीं बताते।
- Crawling से ranking infer नहीं करूंगा। Crawling, indexing, ranking और traffic अलग stages हैं।
curl -I https://example.com/blog/example-post
HTTP/2 308
location: https://example.com/en/blog/example-postयह illustrative example है। Point यह है कि exact status और destination verify किए जाएँ, न कि routing rule के काम करने की assumption रखी जाए।
मैं क्या conclude कर सकता हूँ और क्या नहीं
मैं confirm कर सकता हूँ कि existing 308 redirects ने unexpected /blog/... requests को 404 पर खत्म होने से रोका और intended URLs पर भेजा।
मैं यह confirm नहीं कर सकता कि Yandex sitemap parser इसकी वजह था। मैंने ranking gain, indexing gain या किसी specific amount में “saved” traffic भी measure नहीं किया। ऐसा दावा available evidence से stronger होगा।
मेरे लिए lesson अधिक narrow लेकिन अधिक useful है: URL architecture को system boundaries पर predictable mistakes tolerate करने चाहिए। Clean sitemap पहली defense line है। Precise permanent redirect layer दूसरी।
दोनों मौजूद हों तो unexpected crawler path के हजारों dead URLs में बदलने की संभावना बहुत कम होती है।