Your RAG Is Only as Good as What You Feed It
Crawl4AI is the right tool when the docs you want live behind JavaScript, span hundreds of pages, or come wrapped in so much navigation chrome that your chunks are mostly sidebar. It is the wrong tool when the site has a sitemap and plain static HTML, because requests plus trafilatura does that job in 20 lines and starts instantly.
This post builds the full path a home-labber would run: Crawl4AI crawls a docs site into clean markdown, a script chunks it by heading, Ollama embeds it with nomic-embed-text, Qdrant stores it, and a query script pulls back the closest passages. Everything below was run end to end against Crawl4AI 0.9.4 (released September 2026, Apache-2.0 licensed, as of October 2026) with Qdrant 1.19.1.
Full example: Clone the working files at github.com/KingPin/sumguy-examples/llm/crawl4ai-rag-pipeline
Using a headless browser to read a static page is like hiring a forklift to move a couch. Technically it works, but your neighbors will have questions. So we start with the exit ramp.
Do You Even Need Crawl4AI?
Check these first. Each one is cheaper than a browser fleet.
- The site publishes
llms.txtor a markdown export. Many docs sites do now. Fetchhttps://docs.example.com/llms.txtand you have a curated index of pages, often with.mdlinks. No crawler needed. - The site has an API. GitHub, Discourse, Read the Docs, and most wikis expose content as JSON or raw markdown. Use that.
- Static HTML plus a sitemap. Pull
sitemap.xml, fetch each URL withhttpx, runtrafilatura.extract(). Done. - The docs are a git repo. Clone it. The source markdown is cleaner than anything a crawler produces.
Reach for Crawl4AI when you hit one of these: client-side rendered pages (React or Vue docs that return an empty shell without JavaScript), sites with no sitemap where you must follow links, or pages where you want the page chrome stripped by a relevance filter instead of hand-written selectors.
The Stack: Three Containers and a Venv
Qdrant stores vectors. Ollama runs the embedding model. Crawl4AI runs from a Python venv (the Docker REST server comes later).
services: qdrant: image: qdrant/qdrant:v1.19.1 ports: - "6333:6333" volumes: - qdrant_data:/qdrant/storage restart: unless-stopped
ollama: image: ollama/ollama:latest ports: - "11434:11434" volumes: - ollama_data:/root/.ollama restart: unless-stopped
volumes: qdrant_data: ollama_data:Bring it up and pull the embedding model:
docker compose up -ddocker compose exec ollama ollama pull nomic-embed-textnomic-embed-text produces 768-dimension vectors. I checked: a call to /api/embed returns a list of 768 floats per input. That number matters in a minute, because Qdrant refuses vectors that do not match the collection size. Ollama also has embeddinggemma (300M) if you want something newer. Swap the name and the dimension and nothing else changes.
Now the crawler:
python3 -m venv .venv && . .venv/bin/activatepip install crawl4ai qdrant-client httpxcrawl4ai-setup # downloads the Playwright browserCrawl a Docs Site
Crawl4AI’s deep crawl runs a BFS (breadth-first) strategy over links, with filters that decide which URLs are worth visiting. Here is the whole script:
import asyncioimport hashlibimport jsonimport sysfrom pathlib import Path
from crawl4ai import ( AsyncWebCrawler, BFSDeepCrawlStrategy, BrowserConfig, CacheMode, CrawlerRunConfig, DefaultMarkdownGenerator, PruningContentFilterLXML,)from crawl4ai.deep_crawling import DomainFilter, FilterChain, URLPatternFilter
START_URL = sys.argv[1] if len(sys.argv) > 1 else "https://docs.example.com/"MAX_PAGES = int(sys.argv[2]) if len(sys.argv) > 2 else 50OUT = Path("pages")
async def main() -> None: host = START_URL.split("/")[2] config = CrawlerRunConfig( deep_crawl_strategy=BFSDeepCrawlStrategy( max_depth=2, max_pages=MAX_PAGES, filter_chain=FilterChain( [ DomainFilter(allowed_domains=[host]), URLPatternFilter(patterns=["*/changelog/*", "*/blog/*"], reverse=True), ] ), ), markdown_generator=DefaultMarkdownGenerator( content_filter=PruningContentFilterLXML(threshold=0.48, threshold_type="fixed") ), cache_mode=CacheMode.ENABLED, check_robots_txt=True, semaphore_count=2, mean_delay=1.0, max_range=1.0, verbose=False, user_agent="MyHomelabRAG/1.0 (+https://example.com/bot)", ) OUT.mkdir(exist_ok=True) async with AsyncWebCrawler(config=BrowserConfig(headless=True, verbose=False)) as crawler: results = await crawler.arun(START_URL, config=config) kept = 0 for r in results: if not r.success: print(f"skip {r.url}: {r.error_message}") continue md = r.markdown.fit_markdown or r.markdown.raw_markdown if len(md) < 200: continue name = hashlib.sha1(r.url.encode()).hexdigest()[:12] (OUT / f"{name}.md").write_text(md, encoding="utf-8") (OUT / f"{name}.json").write_text(json.dumps({"url": r.url})) kept += 1 print(f"crawled {len(results)} pages, kept {kept}")
asyncio.run(main())Run it with python crawl.py https://docs.example.com/ 50. A few things in there earn their keep.
max_pages and max_depth are your seatbelt. A docs site with a version switcher can have thousands of URLs. Without a cap, your 2 AM self gets to learn what a 40,000-page crawl does to a Pi.
DomainFilter keeps the crawl on the host. include_external already defaults to off. The filter adds an explicit allow-list: it lets through the host and its subdomains, and blocks sibling domains such as the bare parent domain. The URLPatternFilter with reverse=True excludes matching URLs, which is how you skip changelogs and blog posts that bloat the index with stale version noise.
The result of arun with a deep crawl strategy is a list of results (when stream=False, the default). Each one carries .markdown, an object with raw_markdown, fit_markdown, markdown_with_citations, and references_markdown. It also prints as the raw markdown if you treat it as a string.
Raw Markdown vs Fit Markdown
raw_markdown is the whole page converted to markdown. That includes nav menus, footers, “Edit this page” links, and cookie banners. Embed that and your vector index fills with chunks that all say “Home / Docs / API / Pricing”.
fit_markdown is the same page after a content filter removes low-value blocks. The filter lives on the markdown generator:
PruningContentFilterLXMLscores blocks by text density, link density, and tag weight. Athresholdof 0.48 withthreshold_type="fixed"is the default. No query needed.BM25ContentFilter(user_query="...")keeps blocks that match a query. Use it when you want only the parts of each page about, say, “authentication”.
In 0.9.4 the older PruningContentFilter still works but prints a deprecation warning telling you to switch to the LXML variant, so the code above uses that. This is the kind of change you hit when you copy a 0.4-era tutorial.
On one docs page from a test crawl, the raw markdown was 26,132 characters and the fit version 18,543. That is 29% less text to embed, and the cut text was menus and link lists. Pruning is heuristic, so it sometimes drops a short but useful block. Spot-check a few pages before trusting it with 5,000.
Be Polite or Get Blocked
Crawlers get banned for being rude, and rude is the default here. Check the knobs against what you want:
check_robots_txtdefaults toFalse. Set it toTrue. I tested it against a page that robots.txt disallows: the result came backsuccess=Falsewith status 403 and the message “Access denied by robots.txt”. Crawl4AI caches the parsed rules in a local database, so it does not re-fetch robots.txt for every page.semaphore_countdefaults to 5 concurrent pages. For a docs site you do not own, 2 is plenty.mean_delayandmax_rangeadd a randomized pause between requests. The defaults (0.1 and 0.3 seconds) are for a server you control. I set 1.0 and 1.0 for strangers.user_agentshould say who you are and give a contact URL. A site owner who can reach you rarely needs to block you.
For large batches with arun_many, Crawl4AI also ships a RateLimiter (with exponential backoff on 429 and 503) and a MemoryAdaptiveDispatcher that slows down when RAM gets tight. You pass the dispatcher as arun_many(urls, config=..., dispatcher=...). Deep crawls like the one above use the per-config settings instead.
Robots.txt is a politeness convention, not a legal shield. Terms of service and copyright still apply to what you do with the content. Crawling a project’s own public docs for a private index is low risk, though terms of service vary. Re-publishing someone else’s content is a different story.
JavaScript-Rendered Pages
This is where the browser earns its RAM. For a page that renders content client-side, tell Crawl4AI what “loaded” looks like:
config = CrawlerRunConfig( wait_for="css:.docs-content", # block until this selector exists scan_full_page=True, # scroll to trigger lazy loading delay_before_return_html=0.5,)wait_for accepts a css: selector or a js: expression that returns true. The default wait_until is domcontentloaded, which returns before most single-page apps finish drawing. If your markdown comes back as 40 characters, the page was not ready. Add a wait_for.
Caching: Stop Hammering the Site While You Debug
You will run this script 15 times while tuning chunk sizes. CrawlerRunConfig defaults to CacheMode.BYPASS, which always fetches fresh. The script above uses CacheMode.ENABLED, so the second run reads pages from the local cache and the site sees nothing. The other modes are DISABLED, READ_ONLY, and WRITE_ONLY.
Use ENABLED while iterating. Switch to BYPASS when you want a fresh crawl for real.
Chunk, Embed, Upsert
Crawl4AI ships chunking strategies, but for docs you want headings. A heading is a topic boundary the author already drew for you. Splitting on #, ##, and ### keeps each chunk about one thing, and the heading text goes along with it so the embedding knows the subject.
import jsonimport reimport uuidfrom pathlib import Path
import httpxfrom qdrant_client import QdrantClientfrom qdrant_client.models import Distance, PointStruct, VectorParams
OLLAMA = "http://localhost:11434"MODEL = "nomic-embed-text" # 768 dimensionsCOLLECTION = "docs"MAX_CHARS = 1500
def chunk(md: str) -> list[tuple[str, str]]: parts = re.split(r"(?m)^(?=#{1,3} )", md) out = [] for part in parts: part = part.strip() if len(part) < 80: continue heading = part.splitlines()[0].lstrip("# ").strip() buf = "" for para in part.split("\n\n"): if buf and len(buf) + len(para) > MAX_CHARS: out.append((heading, buf)) buf = f"{heading}\n\n" buf += para + "\n\n" out.append((heading, buf.strip())) return out
def embed(texts: list[str]) -> list[list[float]]: r = httpx.post(f"{OLLAMA}/api/embed", json={"model": MODEL, "input": texts}, timeout=120) r.raise_for_status() return r.json()["embeddings"]
def main() -> None: qd = QdrantClient(url="http://localhost:6333") if not qd.collection_exists(COLLECTION): qd.create_collection(COLLECTION, vectors_config=VectorParams(size=768, distance=Distance.COSINE))
total = 0 for md_file in sorted(Path("pages").glob("*.md")): url = json.loads(md_file.with_suffix(".json").read_text())["url"] chunks = chunk(md_file.read_text(encoding="utf-8")) if not chunks: continue vectors = embed([f"search_document: {text}" for _, text in chunks]) points = [ PointStruct( id=str(uuid.uuid5(uuid.NAMESPACE_URL, f"{url}#{i}")), vector=vec, payload={"url": url, "heading": heading, "text": text}, ) for i, ((heading, text), vec) in enumerate(zip(chunks, vectors)) ] qd.upsert(COLLECTION, points=points) total += len(points) print(f"{url}: {len(points)} chunks") print(f"upserted {total} chunks")
if __name__ == "__main__": main()Three details to know:
- The
search_document:prefix.nomic-embed-textwas trained with task prefixes. Documents getsearch_document:, questions getsearch_query:. The model card asks for them, so skipping them throws away part of what the model learned. Other models (EmbeddingGemma has its own prompt formats) want different prefixes, so read the model card. /api/embed, not/api/embeddings. The current Ollama endpoint takesinputas a string or a list and returnsembeddings, a list of vectors. The older/api/embeddingstakes a singleprompt. Batch with the new one.- Stable IDs. The point ID is a UUID derived from the URL and chunk index. Re-run the ingest and the same points get overwritten, so your index never fills with duplicates. A gotcha: if a page gets shorter, the old trailing chunks stay until you delete them.
On five docs pages from a test crawl, this produced 78 chunks, and a second run left the count unchanged.
Query It
import sys
import httpxfrom qdrant_client import QdrantClient
question = " ".join(sys.argv[1:]) or "How do I limit how fast the crawler hits a site?"r = httpx.post( "http://localhost:11434/api/embed", json={"model": "nomic-embed-text", "input": [f"search_query: {question}"]}, timeout=60,)r.raise_for_status()qd = QdrantClient(url="http://localhost:6333")hits = qd.query_points("docs", query=r.json()["embeddings"][0], limit=3).pointsfor h in hits: print(f"{h.score:.3f} {h.payload['url']} [{h.payload['heading']}]") print(" ", h.payload["text"][:200].replace("\n", " "), "\n")query_points is the current search call. In qdrant-client 1.19.1 the old search method no longer exists, so tutorials that use it fail with an AttributeError.
With only five pages indexed, my results were mediocre. Questions about cache modes found the right quickstart section. A question about robots.txt returned a “How You Can Support” block, because the tiny corpus had nothing better. That is the honest result of a five-page index. Crawl the whole site before you judge retrieval quality. To answer questions, feed the top hits plus the question to a local LLM through the same Ollama server.
The REST API Route
If your ingest code is not Python, or you want the crawler on a different machine from the script, run the official Docker server. As of 0.9.4 the image is unclecode/crawl4ai:0.9.4, it listens on port 11235, wants --shm-size=1g (the README’s run command sets it, and Chromium is known to struggle on Docker’s small default), and requires an API token. Without a token it only answers requests from inside its own container.
export CRAWL4AI_API_TOKEN="$(openssl rand -hex 32)"docker run -d -p 11235:11235 --name crawl4ai --shm-size=1g \ -e CRAWL4AI_API_TOKEN="$CRAWL4AI_API_TOKEN" \ unclecode/crawl4ai:0.9.4
curl -s http://localhost:11235/md \ -H "Authorization: Bearer $CRAWL4AI_API_TOKEN" \ -H "Content-Type: application/json" \ -d '{"url": "https://docs.example.com/", "f": "fit"}'/md returns one page as markdown, and f picks the filter (fit is the pruning one). For full control, POST /crawl takes urls, a browser_config, and a crawler_config, each shaped as {"type": "CrawlerRunConfig", "params": {...}}. I confirmed check_robots_txt and cache_mode pass through that way. The response holds a results list, and each result has the same markdown fields as the Python object.
Pin the tag so a pull next month does not change your API under you.
The SumGuy Take
Run the library for a one-time crawl of a site you care about, and put the REST container on the box if several scripts will share a browser. Turn on check_robots_txt, cap max_pages, cache while you iterate, and use fit markdown. Chunk by heading, keep the URL in the payload so answers can cite their source, and pin your embedding model, because switching models means re-embedding everything.
If the docs are static, skip all of it and use trafilatura. A browser you do not need is a cron job waiting to fail at 3 AM.
Common Questions
Does Crawl4AI respect robots.txt?
Only if you ask. The check_robots_txt option on CrawlerRunConfig defaults to False. Set it to True and Crawl4AI fetches and caches the rules, then returns a 403 result with “Access denied by robots.txt” for disallowed URLs instead of crawling them.
Can Crawl4AI run without a GPU?
Yes. Crawl4AI drives a headless Chromium browser and does not need a GPU. The Docker server wants 1 GB of shared memory (--shm-size=1g), and each open page costs RAM, so keep concurrency low on a small box. A GPU only matters for the embedding or LLM step, and nomic-embed-text runs fine on CPU.
Do I need an LLM to use Crawl4AI?
No. Crawling, markdown conversion, and the pruning and BM25 filters all run without any model. An LLM is optional, for features like LLMExtractionStrategy that pull structured fields out of a page. A RAG ingest pipeline needs only the crawler plus an embedding model.
Which embedding model should I use with Ollama?
Start with nomic-embed-text for a Crawl4AI and Ollama pipeline. The nomic-embed-text model outputs 768-dimension vectors, runs fine on CPU, and is the model this pipeline was tested with. embeddinggemma (300M parameters) is a newer option. Whatever you pick, set the Qdrant collection size to match the model’s dimension, and re-embed everything if you switch.
Is it legal to crawl a docs site for my RAG index?
Often yes for personal use, but this is not legal advice. Crawling a project’s public docs for a private, personal index is low risk. Ignoring robots.txt, hammering servers, or re-publishing copyrighted content is where trouble starts. Read the site’s terms and honor its robots.txt file.