https://github.com/plasmate-labs/plasmate-python
Python SDK for Plasmate - fetch web pages as structured SOM JSON. pip install plasmate.
https://github.com/plasmate-labs/plasmate-python
ai-agents headless-browser llm plasmate python sdk som web-scraping
Last synced: 15 days ago
JSON representation
Python SDK for Plasmate - fetch web pages as structured SOM JSON. pip install plasmate.
- Host: GitHub
- URL: https://github.com/plasmate-labs/plasmate-python
- Owner: plasmate-labs
- License: apache-2.0
- Created: 2026-03-25T12:01:44.000Z (4 months ago)
- Default Branch: master
- Last Pushed: 2026-03-26T18:45:41.000Z (4 months ago)
- Last Synced: 2026-06-08T12:04:59.085Z (about 2 months ago)
- Topics: ai-agents, headless-browser, llm, plasmate, python, sdk, som, web-scraping
- Language: Python
- Homepage: https://docs.plasmate.app/sdk-python
- Size: 13.7 KB
- Stars: 0
- Watchers: 0
- Forks: 0
- Open Issues: 0
-
Metadata Files:
- Readme: README.md
- License: LICENSE
Awesome Lists containing this project
- awesome-plasmate - Python Scraper - Standalone Python scraping toolkit with batch support. (Web Scraping)
README
# Plasmate Python Toolkit
**Simple, powerful web scraping with [Plasmate](https://plasmate.dev)'s Semantic Object Model.**
This toolkit provides a clean Python interface to the Plasmate CLI for developers who want structured web data without the overhead of a full framework like Scrapy.
## The Problem
Traditional scraping with libraries like `requests` and `BeautifulSoup` requires you to write complex, site-specific parsers that break whenever a CSS class or HTML tag changes.
```python
# The old way: brittle, complex, high-maintenance
import requests
from bs4 import BeautifulSoup
url = 'https://news.ycombinator.com/'
response = requests.get(url)
soup = BeautifulSoup(response.text, 'html.parser')
stories = []
for row in soup.select('tr.athing'):
title_el = row.select_one('td.title > span.titleline > a')
stories.append({
'title': title_el.text,
'url': title_el.get('href'),
})
```
**Result**: You get fragile code and raw HTML. Processing this with an LLM is expensive: a typical news homepage can be **10,000-20,000 tokens**.
## The Solution
Plasmate handles fetching and parsing, giving you a clean **Semantic Object Model (SOM)**. This library makes it trivial to use Plasmate in any Python script.
```python
# The new way: simple, robust, low-maintenance
from plasmate_scraper import fetch, extract_links
som = fetch("https://news.ycombinator.com/")
links = extract_links(som)
for link in links:
print(f"{link['text']} -> {link['url']}")
```
**Result**: You get clean data and a massive reduction in token usage for LLM pipelines. The same homepage as a Plasmate SOM is often **1,000-2,000 tokens** — an order of magnitude smaller.
## Installation
```bash
pip install plasmate-scraper
```
You also need the Plasmate CLI:
```bash
# macOS
brew install nicholasgasior/plasmate/plasmate
# From source
cargo install plasmate
# Or download from https://github.com/nicholasgasior/plasmate/releases
```
## Quick Start
### Fetch a single page
```python
from plasmate_scraper import fetch, extract_text
# Fetch the page and get the SOM
som = fetch("https://example.com")
# The SOM is a dictionary
print(som['title'])
# > Example Domain
# Use helpers to extract common data
text = extract_text(som)
print(text)
# > Example Domain
# > This domain is for use in illustrative examples.
# > More information...
```
### Fetch multiple pages concurrently
```python
from plasmate_scraper import batch_fetch
urls = [
"https://news.ycombinator.com",
"https://github.com/explore",
"https://dev.to",
]
# Fetches all pages in parallel
results = batch_fetch(urls, max_concurrent=10)
for som in results:
if 'error' in som:
print(f"Failed to fetch {som['url']}: {som['error']}")
else:
print(f"Fetched: {som.get('title', 'Untitled')}")
```
## API Reference
### `fetch()`
```python
fetch(
url: str,
*,
timeout: int = 30,
javascript: bool = True,
format: str = "json",
binary: str = "plasmate",
extra_args: list[str] | None = None,
) -> dict
```
Fetches a URL and returns the parsed SOM. Raises `PlasmateError` on failure.
### `batch_fetch()`
```python
batch_fetch(
urls: list[str],
*,
max_concurrent: int = 5,
raise_on_error: bool = False,
# ... accepts same args as fetch()
) -> list[dict]
```
Fetches multiple URLs in parallel. If `raise_on_error` is `False` (default), failures are returned as dicts like `{'url': '...', 'error': '...'}`.
### Utility Functions
All utilities take a SOM dictionary as input.
```python
from plasmate_scraper import (
extract_text, # All text content as a string
extract_links, # [{'url': '...', 'text': '...'}]
extract_headings, # [{'level': 1, 'text': '...'}]
extract_tables, # Table regions/elements from the SOM
extract_images, # [{'src': '...', 'alt': '...'}]
extract_by_role, # Filter elements by SOM role
)
```
## Comparison to Alternatives
| Feature | `requests` + `bs4` | `playwright` | `plasmate-scraper` |
|---------|----------------------|--------------|--------------------|
| Parsing | Manual (CSS/XPath) | Manual (CSS/XPath) | **Automatic (SOM)** |
| Resilience | Low (breaks easily) | Low (breaks easily) | **High (semantic)** |
| JS Support | No | Yes | **Yes (default)** |
| Concurrency | Manual (e.g., `ThreadPool`) | Manual | **Built-in (`batch_fetch`)** |
| LLM-ready | No (too verbose) | No (too verbose) | **Yes (token-efficient)** |
## License
Apache 2.0 — see [LICENSE](LICENSE).
---
## Part of the Plasmate Ecosystem
| | |
|---|---|
| **Engine** | [plasmate](https://github.com/plasmate-labs/plasmate) - The browser engine for agents |
| **MCP** | [plasmate-mcp](https://github.com/plasmate-labs/plasmate-mcp) - Claude Code, Cursor, Windsurf |
| **Extension** | [plasmate-extension](https://github.com/plasmate-labs/plasmate-extension) - Chrome cookie export |
| **SDKs** | [Python](https://github.com/plasmate-labs/plasmate-python) / [Node.js](https://github.com/plasmate-labs/quickstart-node) / [Go](https://docs.plasmate.app/sdk-go) / [Rust](https://github.com/plasmate-labs/quickstart-rust) |
| **Frameworks** | [LangChain](https://github.com/langchain-ai/langchain/pull/36208) / [CrewAI](https://github.com/plasmate-labs/crewai-plasmate) / [AutoGen](https://github.com/plasmate-labs/autogen-plasmate) / [Smolagents](https://github.com/plasmate-labs/smolagents-plasmate) |
| **Tools** | [Scrapy](https://github.com/plasmate-labs/scrapy-plasmate) / [Audit](https://github.com/plasmate-labs/plasmate-audit) / [A11y](https://github.com/plasmate-labs/plasmate-a11y) / [GitHub Action](https://github.com/plasmate-labs/som-action) |
| **Resources** | [Awesome Plasmate](https://github.com/plasmate-labs/awesome-plasmate) / [Notebooks](https://github.com/plasmate-labs/notebooks) / [Benchmarks](https://github.com/plasmate-labs/plasmate-benchmarks) |
| **Docs** | [docs.plasmate.app](https://docs.plasmate.app) |
| **W3C** | [Web Content Browser for AI Agents](https://www.w3.org/community/web-content-browser-ai/) |