SiteIndex
Looks at what you already have first, asks only what it cannot see, and turns that into the plan to rank your site or your digital product.
Cuándo se usa
When shipping a new site, preparing a launch, when weeks have passed and it still does not show up, or when you want to rank and do not know where to start.
SKILL.md
# SiteIndex
Make the site discoverable (crawlers get in and understand what each page is),
worth showing (there is a reason to rank it above someone else) and measurable
(it ends in data, not in an opinion).
## Look first, recommend later
Phase 0 runs a sweep against the live domain and reads what the server actually
serves, which is often not what the repository says: whether www and http all end
up at the same address, what robots.txt blocks and which AI bots it names, how
many URLs the sitemap carries and whether they are dated, the whole head, the
JSON-LD types already there, which analytics is installed, how many words the
HTML has before JavaScript runs, and the DNS records that reveal whether the site
is **already verified in Search Console and Bing**.
Only then does it ask, in one go and at most eight things, what the sweep cannot
see: whether anyone opens Search Console and what it says, who the customer is
and what they would type to find you, which pages bring money, who you compete
with, whether anything gets published and what was tried before without success.
Everything goes into a status sheet with three columns (done, missing, not
applicable) and who can fix each item: the agent in the code, or you in a panel
the agent cannot reach. Every line of the report starts with its real status, so
nobody hands you back the work you did a year ago.
## Rule 0: numbers get verified, never recited
Anything with a number in it expires: recommended title lengths, Core Web Vitals
thresholds, required fields for structured data and, above all, bot names. Check
them against the official source at the moment of use. If it cannot be verified,
say it is not verified. Do not invent it.
## So they can get in
**The gatekeepers.** Read the robots.txt that already exists before writing
anything. Block only what is useless or duplicated, especially catalogue filters,
which spawn a URL per combination. It is a sign, not a lock, and blocking there
plus a noindex tag cancels itself out: the bot never gets to read the tag. Never
block the CSS or JavaScript the page needs to render.
**The AI decision, which belongs to the client and not to you.** Two groups, two
different consequences. Training crawlers (GPTBot, ClaudeBot, CCBot,
Google-Extended, Bytespider, Applebot-Extended, meta-externalagent, Amazonbot)
can be blocked without losing visibility. Search and answer crawlers
(OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot,
Perplexity-User) remove you from AI answers if you block them. The classic trap:
Google-Extended does NOT affect your Google ranking, only Gemini training. And the
reverse, which almost nobody knows: blocking it does not keep you out of AI
Overviews either, because those draw from the regular Search index.
**Foundations that are not negotiable.** HTTPS, a single canonical version of the
domain with a 301 from the other, content present in the mobile version, real
404s, and content visible in the served HTML rather than only after JavaScript
runs. If the site has motion, cross over to FrontLaxWeb here: scroll storytelling
is where that failure is born, because the copy ends up inside something that only
exists once JavaScript mounts.
## So they understand each page
**The head, page by page.** Its own title and description, a single h1, a
canonical when the same content is reachable through several URLs, alt text on
images, and Open Graph with an image, which is what shows when the link is pasted
in a chat.
**Several languages.** One URL per language (/es/, /en/): that slash you see in so
many addresses is not detection, it is that they are separate pages. They are
annotated with hreflang, and the three rules that break most often are that each
version must list itself, that the links must go both ways, and that there must be
an x-default. Each language canonical points at ITSELF: if the one on /es/ points
to /en/, the Spanish version has just been removed from the index. And no
automatic redirect by language or country, which is the big one: Googlebot crawls
without an Accept-Language header and mostly from US addresses, so it lands on the
same version every time and the rest never get indexed.
**JSON-LD.** It must describe what is VISIBLE on the page, and be validated before
it is trusted. That said, Google states in writing that structured data is not
required to appear in its AI features, so it is not sold as a trick.
**Architecture and internal links.** What earns money sits a few clicks from the
homepage. No orphan pages, because a sitemap is not a substitute for a link.
Descriptive anchor text, real links instead of JavaScript buttons, and topic
clusters linked both ways.
## So it is worth showing
**Content, which is what actually ranks.** Pull the real questions (from
customers, from the site search, from Search Console), sort them by intent (know,
compare, go, buy), and give each intent ONE page. Answer at the top, in the first
two or three sentences. Sign it and date it honestly. And revise the old page
before writing a new one, which usually pays better. With Google's own
self-assessment questions in front of you, and knowing E-E-A-T is not a tag you
switch on.
**Speed.** LCP, INP and CLS, measured with real visitors rather than a simulation.
Images, fonts, JavaScript and server, in that order. Speed will not save content
that does not answer the query, but it decides between two similar pages.
**Digital product, if what you sell or give away is software.** Nobody searches for your
name until they already know you, so the homepage has to win on category and problem
rather than on the brand. The pages that bring users are use cases, the comparison, the
"alternative to X", pricing, download, indexable documentation and a changelog with real
dates. An assistant only recommends you if it can read on the page what it is, which
system it runs on, what it costs, under which licence, what data it touches and how to
install it. And half the discovery happens off your own site: package managers,
alternative directories, launch platforms and GitHub. Careful with spinning a hundred
"alternative to" pages off one template: that is scaled content abuse and doorway pages at
the same time.
**Local business, if you serve an area.** Google says local results come down to
relevance, distance and prominence. Distance cannot be changed; a complete
verified profile, consistent details across the site and answered reviews can. And
you cannot pay for better local ranking: Google says so in those words.
## So people know, and so you can check
**Registration and pings.** Search Console as a domain property, Bing Webmaster
Tools as the second opinion, and IndexNow to announce the moment you publish.
Google does not take part in IndexNow, so it complements rather than replaces.
**Getting cited by AI.** Google published its official guidance in May 2026 and it
is blunt: there is no separate algorithm, its AI features run on the same Search
systems, and you do not need special files, markup or Markdown to appear. What
does control what they can use from your page are the snippet directives, and
cutting those cuts your own presence. AEO and GEO, done properly, are SEO.
**Measure, which is where the work ends.** Search Console comparing two periods,
the queries where you already rank low (the most profitable to-do list there is),
PageSpeed, server logs, and asking the assistants your own key questions once a
month to see whether they cite you and who they cite alongside. With the date
written down, or there is nothing to compare against later.
## The tips that pay off most
Start with the queries already sitting on page two, where the least work turns into the
most visits. Revise before you publish. Write titles that say what the person gains, not
your internal product name. Answer in the first three sentences. Link from your strong
pages to the ones you want to lift. Put the price in text. Fewer and better pages, because
the weak ones drag the good ones down. Check what the bot sees rather than what you see.
Sign it, date it and say how you made it. Turn the questions that reach your inbox into
pages. Announce new content instead of waiting for the crawl. If there is video, publish
the transcript. And repeat the same measurements every month, with the date written down.
## The failures that break indexing most often
Before touching anything on a site that does not show up, rule these out in order:
a noindex left over from development, a site-wide Disallow inherited from staging,
a new site never registered and with no inbound links, both domain versions live
at once, a canonical pointing at another page, content that only exists after
JavaScript runs, an automatic language redirect, orphan pages, and the most common
case on perfectly built sites: indexed, with no reason to be shown.
## What is out of scope
Writing the content, paid advertising, email deliverability, and full audits of
very large sites that need a paid crawler. Say so plainly instead of improvising,
and for email hand over the tools: Spamhaus and MXToolbox Email Health, both in
the resources section of this site.