# SuperScraper — Agent Skill

SuperScraper is a web scraping + structured-extraction API for AI agents: turn any URL (or a search query) into clean markdown or typed JSON, no browsers or proxies to manage. Firecrawl-compatible request shapes; page scrapes report how each page was fetched, and structured data, when found, carries a completeness score so an agent knows what to trust.

Deployed host (base URL for all examples below):

```
https://api.superscraper.dev
```

---

## Step 1 — Get an API key

`POST /v1/keys/provision` is public and mints your first key. Do this once, then reuse the key.

```bash
curl -s -X POST https://api.superscraper.dev/v1/keys/provision \
  -H "Content-Type: application/json" \
  -d '{"email":"you@example.com"}'
# => { "apiKey": "ss_live_...", "tenantId": "..." }
```

Send the key on every `/v1/*` call as a Bearer token (or the `x-api-key` header):

```
Authorization: Bearer ss_live_...
```

---

## Step 2 — Core calls

All bodies are JSON. Formats are Firecrawl-compatible: `{ "url": "...", "formats": ["markdown","json"] }`.

### Scrape a URL → markdown / json / links / screenshot

```bash
curl -s -X POST https://api.superscraper.dev/v1/scrape \
  -H "Authorization: Bearer $SS_KEY" -H "Content-Type: application/json" \
  -d '{"url":"https://example.com","formats":["markdown","links"]}'
```

### Extract typed data with a schema

```bash
curl -s -X POST https://api.superscraper.dev/v1/extract \
  -H "Authorization: Bearer $SS_KEY" -H "Content-Type: application/json" \
  -d '{"url":"https://example.com/product","schema":{"name":"string","price":"number","inStock":"boolean"}}'
```

(No `schema` => free deterministic JSON-LD cascade, no LLM cost.)

### Search the web (ground your answers)

```bash
curl -s -X POST https://api.superscraper.dev/v1/search \
  -H "Authorization: Bearer $SS_KEY" -H "Content-Type: application/json" \
  -d '{"query":"best roofing contractors in Austin TX","limit":10}'
```

Other core endpoints (same auth header, JSON body):

| Endpoint | Purpose |
| --- | --- |
| `POST /v1/scrape` | URL → markdown/json/links/screenshot (`formats[]`) |
| `POST /v1/extract` | Schema-driven structured extraction |
| `POST /v1/map` | Discover all URLs on a site (`{ "url": "..." }`) |
| `POST /v1/crawl` | Async multi-page crawl → `jobId`; poll `GET /v1/jobs/:id` |
| `POST /v1/search` | Web search (cascades SERP → DuckDuckGo → Bing) |
| `POST /v1/batch` | Scrape up to 100 URLs in one request |
| `POST /v1/parse` | PDF/DOCX (upload or `{ "url" }`) → markdown; `"format":"rag"` → embeddings-ready chunks (content-hash cached) |
| `POST /v1/enrich` | Domain → emails/phones/social/tech/team with per-field provenance |

---

## Step 3 — Hosted MCP is not available

Do not configure a hosted MCP URL. Use the REST API in Step 2.
Local stdio MCP is developer-only (source checkout) and is not the public hosted surface.

---

## Step 4 — CLI

```bash
npx superscraper init        # save your key + auto-wire MCP into your IDE (--all)
npx superscraper scrape https://example.com
npx superscraper extract https://example.com/product --schema ./schema.json
npx superscraper pull "plumbers in Austin TX"
```

---

## Step 5 — Machine-readable index

- `GET /llms.txt` — condensed capability index (text/plain)
- `GET /llms-full.txt` — full reference with request/response examples
- `GET /openapi.json` — OpenAPI 3.1 spec (feed to any OpenAPI tool / GPT Action / Composio)
- `GET /agent-onboarding/MANIFEST.md` — the full agent skill manifest (every endpoint, in depth)

```bash
curl -s https://api.superscraper.dev/llms.txt
curl -s https://api.superscraper.dev/openapi.json
```

---

## Conventions you can rely on

- Every error returns `{ "error": "..." }` with a 4xx/5xx status — endpoints never throw raw.
- Scrape responses carry `metadata.fetchMethod`. Structured results (a scrape listing, a no-schema extract) carry `_completeness` (0–1, how complete the result is) and `_extraction_method`; extract adds `_provenance`. `/v1/resolve` carries `_match_score` (how the company was matched) instead. Search and parse results carry no score.
- Default extraction is deterministic (JSON-LD) and free; LLM cost only when you pass a `schema` or use enrichment endpoints.
