Last updated: 3 August 2026
How Does Perplexity Choose Its Sources? A UK Business Guide to AI Search Citations
Perplexity chooses its sources through a retrieval-augmented generation process: its crawler pulls roughly 5–10 candidate pages per query, then a ranking pipeline filters this down to just 3–4 sources that actually get cited, based on relevance, authority, freshness, and corroboration across multiple pages (AuthorityTech.io, 2026).
Key Takeaways
- Perplexity retrieves 5–10 candidate pages per query but only cites 3–4 of them in its final answer, according to AuthorityTech.io (2026).
- Reddit accounts for 20–24% of all Perplexity citations, the highest single-domain concentration of any major AI search engine, per Everything-PR Research (2026).
- Perplexity's median citation density is 6.4 unique domains per answer, more than double ChatGPT's 3.1, according to Attrifast (2026).
- Wikipedia is the single most cited source across Perplexity's search results, with Stack Overflow, GitHub and Medium dominating technical queries, per NP Digital (2026).
- Perplexity's own crawler, PerplexityBot, is documented to respect robots.txt disallow directives, meaning sites that block it will not have their content indexed or cited, according to Perplexity's official Help Center.
How Does Perplexity's Source Selection Process Actually Work?
Perplexity operates on a retrieval-augmented generation (RAG) model, meaning it doesn't rely purely on a pre-trained language model's memory — it actively searches the live web, retrieves candidate documents, and then generates an answer grounded in what it finds. When a user submits a query, Perplexity's crawler fetches a shortlist of pages, typically around 10 pages per query, but only 3 to 4 of them survive into the cited response (AI Labs Audit, 2026).
This is a fundamentally different model to Google's search index, which returns ranked lists of links for the user to click through themselves. Perplexity instead does the reading for the user, synthesising an answer and showing only the sources it judged reliable enough to underpin that synthesis. Co-founder and CEO Aravind Srinivas has described the product succinctly: "Perplexity is a marriage of Wikipedia and ChatGPT."
The retrieval-to-citation funnel
The gap between pages retrieved and pages cited is the single most important number for any UK business trying to understand Perplexity. Being crawled is not the same as being cited — your page has to survive a multi-stage filtering process after retrieval.
| Stage | What happens | Approximate volume |
|---|---|---|
| Initial retrieval | Crawler pulls candidate pages matching query intent | 5–10 pages (AuthorityTech.io, 2026) / ~10 pages (AI Labs Audit, 2026) |
| Ranking pipeline | Pages scored on relevance, authority, freshness | Narrows the pool |
| Final citation | Sources actually referenced in the generated answer | 3–4 pages |
What Ranking Signals Determine Which Sources Perplexity Cites?
Perplexity's ranking pipeline weighs multiple signals when deciding which of its retrieved candidates make the final cut, including topical relevance to the query, domain authority, content freshness, and corroboration — where the same fact appears across several independent sources. A July 2026 study covering 366,000 citations found Perplexity shows a clear tier-1 journalism source preference when answering news-adjacent queries (Kai-Cheng Yang arXiv study via Everything-PR, 2026). No official, fully itemised weighting has been published by Perplexity itself, so specific percentage breakdowns circulating in SEO blogs should be treated as third-party estimates rather than confirmed algorithmic detail.
Relevance and semantic matching
Perplexity's system needs to match the intent behind a query, not just the keywords in it. A page that directly and clearly answers a specific question — in plain, well-structured prose — is easier for the retrieval system to match than a page buried in marketing language. This is precisely why structuring content around direct questions and self-contained answers, the way Aether Agency Ltd builds content for clients, tends to perform well.
Authority and trust signals
Domain-level trust matters. Perplexity leans heavily on established reference sites and recognised publishers for many query types, which is why encyclopaedic and journalistic sources dominate broad informational queries.
Freshness
For time-sensitive queries — pricing, news, regulatory changes — Perplexity favours recently updated content. A page last updated in 2022 is a weaker candidate for a "2026" query than one demonstrably refreshed this year, particularly for UK regulatory or pricing topics where guidance changes annually.
Corroboration across multiple sources
Where several independent, credible pages state the same fact, Perplexity's synthesis process treats that agreement as a trust signal. This is one reason a business's own website alone is rarely enough — being mentioned consistently across trade press, review platforms, and forums strengthens the case for citation.
Which Websites Does Perplexity Cite Most Often?
Perplexity's citation patterns differ sharply from a traditional Google results page. Wikipedia is the most cited source overall across Perplexity's search responses, with Stack Overflow, GitHub, and Medium ranking highly specifically for technical and developer-focused queries (NP Digital, 2026). Community platforms punch well above their weight too: Reddit accounts for 20–24% of all Perplexity citations, the highest single-domain concentration recorded across any major AI search engine, with Perplexity citing an average of 8.2 sources per answer — 3.4 times the citation density of ChatGPT (Everything-PR Research, 2026).
Professional and B2B platforms also feature, though less prominently. A cross-engine study of 325,000 prompts by SEMrush found LinkedIn cited in 5.3% of Perplexity responses, and notably, 59% of those LinkedIn citations came from Company Pages rather than individual posts (Contently, citing SEMrush, 2026). For UK businesses, this is a practical signal: a well-maintained LinkedIn Company Page carries more weight in Perplexity's eyes than founder or employee personal posts.
Comparing Perplexity's citation behaviour to other AI engines
Perplexity is markedly more citation-dense than its major competitors. A 1,200-prompt benchmark spanning 12 verticals found:
| AI Engine | Median unique domains cited per answer |
|---|---|
| Perplexity | 6.4 |
| Claude | 3.6 |
| ChatGPT | 3.1 |
| Gemini | 2.4 |
Perplexity cites 2.1 times more URLs per answer than ChatGPT (Attrifast, 2026). For a UK business, this means Perplexity presents proportionally more opportunities to be cited than most rival engines — but also a more competitive, community- and journalism-heavy field to break into, given Reddit and tier-1 press dominate so many query types.
Does Perplexity Respect Robots.txt When Crawling Websites?
Yes — according to Perplexity's own documentation, PerplexityBot is designed to respect robots.txt disallow directives, meaning it will not index full or partial page text from sites that explicitly block it (Perplexity Help Center). Perplexity maintains published crawler documentation listing its user agents — PerplexityBot for indexing and Perplexity-User for real-time, user-triggered fetches — along with IP address endpoints webmasters can use to configure firewalls (Perplexity Crawlers documentation).
This is not, however, an uncontested area. Independent technical analysis has scrutinised PerplexityBot's user-agent behaviour and crawl classification in detail (51Degrees), and there has been public reporting of tension between Perplexity and publishers over crawling practices, including a Wired investigation that CEO Aravind Srinivas responded to directly. Srinivas stated: "Perplexity is not ignoring the Robot Exclusions Protocol and then lying about it," while also noting that robots.txt is "not a legal framework" but rather a voluntary web standard.
The undeclared crawler question
Separately, reporting has summarised a Cloudflare report describing undeclared or rotating Perplexity-linked crawlers that appeared to bypass robots.txt blocks on some sites (Soar Agency summary). This matters for UK businesses deciding whether to block or welcome AI crawlers: the technical picture is more contested than Perplexity's official documentation alone suggests, and site owners should monitor server logs rather than assume robots.txt directives are the final word.
PerplexityBot vs Perplexity-User: What's the Difference?
Perplexity operates two distinct crawler agents, and understanding the difference matters for any UK website owner configuring access controls. PerplexityBot is the indexing crawler that builds Perplexity's background knowledge base over time, generally respecting robots.txt disallow rules. Perplexity-User is a separate, real-time agent that fetches a specific page the moment a user's query requires it — for example, when someone asks Perplexity to summarise a particular article they've linked to. Both agents are documented with their own user-agent strings and IP ranges in Perplexity's official crawler documentation, and webmasters can configure robots.txt and web application firewalls (WAFs) differently for each.
Comparison: PerplexityBot vs Perplexity-User
| Feature | PerplexityBot | Perplexity-User |
|---|---|---|
| Purpose | Background indexing for future queries | Real-time fetch triggered by a live user query |
| Robots.txt compliance | Documented to respect disallow directives | Behaviour may differ; triggered on-demand |
| Typical trigger | Scheduled or periodic crawling | Immediate, query-specific request |
| Business relevance | Affects long-term citation eligibility | Affects whether a specific page can be summarised live |
How Can UK Businesses Get Cited by Perplexity?
Getting cited by Perplexity is not about traditional keyword stuffing — it requires content structured for extraction, hosted on a crawlable site, and corroborated elsewhere on the web. Given that only 3–4 of every 5–10 retrieved pages survive into a final answer, the margin for winning a citation is genuinely tight, and businesses need a deliberate strategy rather than an incidental one.
Structure content to answer questions directly
Perplexity's retrieval system favours content that opens with a clear, self-contained answer before elaborating. Long preambles, vague headlines, and buried conclusions all reduce the chance that a passage is extracted cleanly. This is the core discipline Aether Agency Ltd applies when building content for clients targeting AI search visibility, not just Google rankings.
Maintain an active LinkedIn Company Page
Given that 59% of Perplexity's LinkedIn citations pull from Company Pages rather than personal posts (Contently, citing SEMrush, 2026), UK B2B firms should treat their Company Page as a genuine citation asset, not an afterthought — kept current with services, credentials, and clear descriptions of what the business does.
Don't ignore community platforms
With Reddit representing 20–24% of all Perplexity citations (Everything-PR Research, 2026), a presence — or at least a favourable reputation — in relevant subreddits and community discussions can meaningfully influence whether a brand surfaces in Perplexity's answers.
Ensure technical crawlability
Confirm robots.txt does not inadvertently block PerplexityBot if visibility in Perplexity is a goal, and check server logs for both PerplexityBot and Perplexity-User activity using the IP ranges published in Perplexity's crawler documentation.
In-House SEO vs Specialist AI Search Support: Which Approach Wins Citations?
Many UK businesses already run in-house SEO efforts focused on Google rankings, and assume the same content will naturally perform in Perplexity. In practice, the two systems reward different things: Google ranks pages, Perplexity extracts and synthesises passages, favouring different structural and authority signals along the way. Businesses relying solely on legacy SEO tactics — keyword density, backlink volume, meta tag optimisation — often find their content is retrieved by Perplexity's crawler but fails to survive the ranking pipeline into an actual citation.
A specialist approach, by contrast, treats AI search citation as a distinct discipline: structuring passages for extraction, building corroborating mentions across community and press platforms, and monitoring crawler behaviour directly. For businesses without the internal resource to run both disciplines in parallel, working with a specialist studio that understands both traditional SEO and AI search citation mechanics — such as Aether Agency Ltd — is typically the faster route to visibility across Google, ChatGPT, and Perplexity simultaneously.
Your Perplexity Source Optimisation Checklist
- Check your robots.txt file does not block PerplexityBot if AI search visibility is a priority for your business.
- Review server logs for PerplexityBot and Perplexity-User activity using the IP ranges in Perplexity's official documentation.
- Restructure key pages so the first 130–160 words fully answer the page's core question in self-contained prose.
- Keep your LinkedIn Company Page complete, current, and detailed — not just your personal profile.
- Monitor relevant Reddit communities and industry forums where your brand or sector is discussed.
- Refresh time-sensitive content (pricing, regulations, statistics) at least annually to preserve freshness signals.
- Seek corroborating mentions across trade press, review sites, and community platforms, not just your own domain.
- Audit which competitors are currently being cited by Perplexity for your priority queries and identify the gap.
FAQ
How does Perplexity choose which websites to cite in its answers?
Perplexity retrieves a shortlist of candidate pages for each query — typically 5 to 10 — then runs them through a ranking pipeline that scores relevance, domain authority, content freshness, and corroboration across sources, ultimately citing only 3 to 4 pages in the final answer (AuthorityTech.io, 2026).
What is Retrieval-Augmented Generation (RAG) and how does Perplexity use it?
Retrieval-Augmented Generation is a technique where an AI system searches the live web for relevant documents before generating its answer, rather than relying solely on pre-trained knowledge. Perplexity uses RAG to fetch current web pages, then synthesises a cited answer grounded in that retrieved content, which is why CEO Aravind Srinivas has described the product as combining the structured referencing of Wikipedia with the conversational fluency of ChatGPT.
Does Perplexity respect robots.txt when crawling websites?
Yes, according to Perplexity's own Help Center documentation, PerplexityBot is designed to respect robots.txt disallow directives and will not index text from blocked sites. However, independent reporting has raised questions about undeclared crawler behaviour, so the picture isn't entirely settled (Soar Agency).
What is the difference between PerplexityBot and Perplexity-User?
PerplexityBot is the background indexing crawler that builds Perplexity's knowledge base over time, while Perplexity-User is triggered in real time whenever a live user query requires fetching a specific page. Both are documented separately in Perplexity's crawler documentation, with distinct IP ranges for webmasters to configure.
How many sources does Perplexity typically cite per answer?
Perplexity's median citation density is 6.4 unique domains per answer, more than double ChatGPT's 3.1 and significantly higher than Claude's 3.6 or Gemini's 2.4 (Attrifast, 2026). Separately, research shows Perplexity averages 8.2 sources per answer overall, 3.4 times ChatGPT's density (Everything-PR Research, 2026).
Why does Perplexity cite Reddit and social media so heavily?
Reddit accounts for 20–24% of all Perplexity citations, the highest single-domain concentration of any major AI search engine (Everything-PR Research, 2026). This likely reflects Perplexity's weighting toward recent, corroborated community discussion for many query types, particularly opinion-based or experience-based questions where forum discussion offers direct, first-hand answers.
How can businesses get their content cited by Perplexity?
Businesses improve their citation chances by structuring content to answer specific questions directly within the first 130–160 words, ensuring their site is crawlable by PerplexityBot, maintaining an active LinkedIn Company Page, and building corroborating mentions across trade press and community platforms like Reddit. Specialist support from an agency experienced in AI search optimisation, such as Aether Agency Ltd, can accelerate this process considerably.
Winning Perplexity Citations with Aether Agency Ltd
Understanding how Perplexity chooses its sources is only useful if a business can act on it — and that's precisely where Aether Agency Ltd's work sits, building content and web presences engineered for the retrieval-to-citation funnel this article describes, not just for traditional Google rankings. As a full-service creative studio covering brand identity, website development, and marketing, Aether Agency Ltd approaches AI search visibility as a technical and editorial discipline in its own right, structuring client content so it survives Perplexity's ranking pipeline rather than simply being crawled and discarded.
Aether Agency Ltd's core proposition is straightforward: brand identity, website development, and marketing that gets clients found on Google, ChatGPT, and Perplexity — the three engines that increasingly shape how UK buyers discover and evaluate businesses. That means every project considers not just how a page ranks, but whether its structure, freshness, and cross-platform corroboration give it a genuine chance of being one of the 3–4 sources Perplexity actually cites.
If your business wants to understand where it currently stands in Perplexity's answers — and what it would take to close the gap — get in touch with Aether Agency Ltd for a conversation about your website, your content, and your visibility across AI search engines.
Related Reading
- AI Search Agency vs Traditional SEO Agency: 2026 UK Guide
- AI Search Readiness Checklist 2026 | Aether Agency Ltd
- B2B Lead Generation from AI Search 2026 | Aether Agency
See How Your Brand Appears in AI Search
Aether AI monitors your visibility across ChatGPT, Perplexity, Google AI Overviews, and Claude in real time. Find out where you stand and what to fix.
Explore Aether AI