Datacrawl quickstart
Three endpoints, one credit meter: POST /v1/search (web/news
search), POST /v1/fetch (block-aware page fetch, optional JS
rendering), POST /v1/extract (clean article text, JSON-LD, or
rule-based fields). Authenticate every call with
Authorization: Bearer <your key> — get a key at
/portal/signup.
Prefer to click first? Signed-in customers can run the same
/v1/extract call from the browser at
/portal/crawl — one field, one button, billed to
the same key.
See also: credits & pricing · auth & errors · OpenAPI 3.1
curl
curl -s https://datacrawl.dev/v1/search \
-H "Authorization: Bearer ak_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"q": "latest EU AI act obligations", "kind": "news", "count": 5}'
curl -s https://datacrawl.dev/v1/fetch \
-H "Authorization: Bearer ak_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com/", "render": false}'
curl -s https://datacrawl.dev/v1/extract \
-H "Authorization: Bearer ak_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com/article", "mode": "readability"}'
Python
import httpx
BASE = "https://datacrawl.dev"
HEADERS = {"Authorization": "Bearer ak_YOUR_KEY"}
hits = httpx.post(f"{BASE}/v1/search", headers=HEADERS,
json={"q": "latest EU AI act obligations", "count": 5},
timeout=60)
hits.raise_for_status()
page = httpx.post(f"{BASE}/v1/fetch", headers=HEADERS,
json={"url": hits.json()["results"][0]["target_url"]},
timeout=120)
page.raise_for_status()
# Extract from the HTML we already fetched — don't pass "url" here, that
# would re-fetch (and re-charge for) the same page. html-input extract is
# just the flat 1-credit surcharge (no fetch cost), vs. url-input extract
# which also charges for the fetch tier used.
text = httpx.post(f"{BASE}/v1/extract", headers=HEADERS,
json={"html": page.json()["html"],
"mode": "readability"}, timeout=120)
print(text.json()["content"])
print("credits left:", text.headers["X-Credits-Remaining"])
TypeScript
const BASE = "https://datacrawl.dev";
const headers = {
Authorization: `Bearer ${process.env.DATACRAWL_API_KEY}`,
"Content-Type": "application/json",
};
const search = await fetch(`${BASE}/v1/search`, {
method: "POST", headers,
body: JSON.stringify({ q: "latest EU AI act obligations", count: 5 }),
});
const { results } = await search.json();
const extract = await fetch(`${BASE}/v1/extract`, {
method: "POST", headers,
body: JSON.stringify({ url: results[0].target_url, mode: "readability" }),
});
console.log((await extract.json()).content);
console.log("credits left:", extract.headers.get("X-Credits-Remaining"));
Good to know
render: trueruns a real browser for JavaScript-heavy pages (costs more — see pricing). If a plain fetch gets blocked and we escalate to the browser for you, the cost difference is charged after the fact./v1/extractaccepts eitherurl(we fetch it) orhtml(you already have it — cheapest).- We honor the target's
robots.txtfor thedatacrawluser-agent. A disallowed path answers403, naming the site's own robots URL, and costs nothing. Send"respect_robots": falseto take that responsibility yourself, as the Terms require. - You pay for content, not attempts: if the target blocks every route we
may use, the transport fails, or you stop waiting, the up-front charge is
refunded and the response carries
X-Credits-Refunded. A page we did fetch but that yields no readable text is still billed. - Honest capacity note: search rides upstream provider quota — bursts may be briefly throttled even below your plan's rate limit, and fetch latency depends on the target site (rendered fetches take seconds, not milliseconds).