{"id":25750419,"url":"https://github.com/antoinejeannot/daidai","last_synced_at":"2025-05-12T16:29:54.291Z","repository":{"id":278758522,"uuid":"935988091","full_name":"antoinejeannot/daidai","owner":"antoinejeannot","description":"Modern dependency \u0026 assets management library ","archived":false,"fork":false,"pushed_at":"2025-03-11T11:44:34.000Z","size":3412,"stargazers_count":5,"open_issues_count":0,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-04-12T05:20:16.119Z","etag":null,"topics":["ai","aiops","dependency-injection-library","llmops","ml","mlops","python"],"latest_commit_sha":null,"homepage":"https://antoinejeannot.github.io/daidai/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/antoinejeannot.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":"SECURITY.md","support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2025-02-20T10:52:02.000Z","updated_at":"2025-03-11T11:44:07.000Z","dependencies_parsed_at":"2025-03-11T12:20:43.185Z","dependency_job_id":"3967b749-6e5f-41c4-963e-6710f252f187","html_url":"https://github.com/antoinejeannot/daidai","commit_stats":null,"previous_names":["antoinejeannot/daidai"],"tags_count":7,"template":false,"template_full_name":"antoinejeannot/cookiecutter-python","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/antoinejeannot%2Fdaidai","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/antoinejeannot%2Fdaidai/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/antoinejeannot%2Fdaidai/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/antoinejeannot%2Fdaidai/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/antoinejeannot","download_url":"https://codeload.github.com/antoinejeannot/daidai/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":253776501,"owners_count":21962499,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["ai","aiops","dependency-injection-library","llmops","ml","mlops","python"],"created_at":"2025-02-26T13:16:39.647Z","updated_at":"2025-05-12T16:29:54.277Z","avatar_url":"https://github.com/antoinejeannot.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cp align=\"center\"\u003e\n    \u003cimg src=\"https://raw.githubusercontent.com/antoinejeannot/daidai/assets/logo.svg\" alt=\"daidai logo\" width=\"200px\"\u003e\n\u003c/p\u003e\n\u003ch1 align=\"center\"\u003e daidai 🍊\u003c/h1\u003e\n\u003cp align=\"center\"\u003e\n  \u003cem\u003eModern dependency \u0026 assets management library for MLOps\u003c/em\u003e\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n\u003ca href=\"https://github.com/antoinejeannot/daidai/actions/workflows/tests.yml\"\u003e\u003cimg src=\"https://github.com/antoinejeannot/daidai/actions/workflows/tests.yml/badge.svg\" alt=\"Tests\"\u003e\u003c/a\u003e\n\u003ca href=\"https://pypi.org/project/daidai/\"\u003e\u003cimg src=\"https://img.shields.io/pypi/v/daidai.svg\" alt=\"PyPI version\"\u003e\u003c/a\u003e\n\u003ca href=\"https://pypi.org/project/daidai/\"\u003e\u003cimg src=\"https://img.shields.io/pypi/pyversions/daidai.svg\" alt=\"Python Versions\"\u003e\u003c/a\u003e\n\u003ca href=\"https://github.com/antoinejeannot/daidai/blob/main/LICENSE\"\u003e\u003cimg src=\"https://img.shields.io/github/license/antoinejeannot/daidai.svg\" alt=\"License\"\u003e\u003c/a\u003e\n\u003cimg alt=\"AI MLOps\" src=\"https://img.shields.io/badge/AI-MLOps-purple\"\u003e\n\u003c/p\u003e\n\n**daidai 🍊** is a minimalist, type-safe dependency management system for AI/ML components that streamlines workflow development with dependency injection, intelligent caching and seamless file handling.\n\n\u003e [!WARNING]\n\u003e **daidai** is still very much a work in progress and is definitely not prod-ready: API has been prioritized over testing, and is still subject to change.\n\u003e\n\u003e It is currently developed as a _[selfish software](https://every.to/source-code/selfish-software)_ to become my personal go-to MLOps library, as well as to distill years of building and maintaining complex NLP pipelines, but feel free to give it a try :)\n\n\u003cdetails open\u003e\n\u003csummary\u003e🎧 Stop reading, listen to daidai 🍊's podcast instead 🎙️\u003c/summary\u003e\n\n[![YouTube](http://i.ytimg.com/vi/ZFaAxzrvucc/hqdefault.jpg)](https://www.youtube.com/watch?v=ZFaAxzrvucc)\n\n_Powered by [NotebookLM](https://notebooklm.google.com/notebook/e273780a-e0db-457f-bb5b-cca0778abe9d/audio)_\n\u003c/details\u003e\n\n## Why daidai?\n\nBuilt for both rapid prototyping and production ML workflows, **daidai 🍊**:\n\n- 🚀 **Accelerates Development** - Reduces iteration cycles with zero-config caching\n- 🧩 **Simplifies Architecture** - Define reusable components with clear dependencies\n- 🔌 **Works Anywhere** - Seamless integration with cloud/local storage via fsspec (local, s3, gcs, az, ftp, hf..)\n- 🧠 **Stays Out of Your Way** - Type-hint based DI means minimal boilerplate\n- 🧹 **Manages Resources** - Automatic cleanup prevents leaks and wasted compute\n- 🧪 **Enables Testing** - Inject mock dependencies / stubs with ease for robust unit testing\n- 🪶 **Requires Zero Dependencies** - _Zero-dependency Core_ philosophy, install optionals at will\n- λ **Promotes Functional Thinking** - Embraces pure functions, immutability, and composition for predictable workflows\n- 🧰 **Adapts to Your Design** - pure functions enable seamless integration with your preferred caching, versioning, validation systems..\n\n\n\u003e **daidai** is named after the Japanese word for \"orange\" 🍊, a fruit that is both sweet and sour, just like the experience of managing dependencies in ML projects. \u003cbr/\u003eIt is being developed with user happiness in mind, while providing great flexibility and minimal boilerplate. It has been inspired by [pytest](https://github.com/pytest-dev/pytest), [modelkit](https://github.com/Cornerstone-OnDemand/modelkit), [dependency injection \u0026 testing](https://antoinejeannot.github.io/nuggets/dependency_injection_and_testing.html) principles and functional programming.\n\n\n## Installation\n\n```bash\n# Core functionality of assets \u0026 predictors\npip install daidai\n\n# Full installation with all features: artifacts, memory tracking, CLI\npip install daidai[all] # or any combination of [artifacts, memory, cli]\n```\n\n## Quick Start\n\n```python\nimport base64\nfrom typing import Annotated, Any\n\nimport openai\n\nfrom daidai import ModelManager, asset, predictor\n\n# Define assets which are long-lived objects\n# that can be used by multiple predictors, or other assets\n@asset\ndef openai_client(**configuration: dict[str, Any]) -\u003e openai.OpenAI:\n    return openai.OpenAI(**configuration)\n\n# Fetch a distant file from HTTPS, but it can be from any source: local, S3, GCS, Azure, FTP, HF Hub, etc.\n@asset\ndef dogo_picture(\n    picture: Annotated[\n        bytes,\n        \"https://images.pexels.com/photos/220938/pexels-photo-220938.jpeg\",\n        {\"cache_strategy\": \"no_cache\"},\n    ],\n) -\u003e str:\n    return base64.b64encode(picture).decode(\"utf-8\")\n\n\n# Define a predictor that depends on the previous assets\n# which are automatically loaded and passed as an argument\n@predictor\ndef ask(\n    message: str,\n    dogo_picture: Annotated[str, dogo_picture],\n    client: Annotated[openai.OpenAI, openai_client, {\"timeout\": 5}],\n    model: str = \"gpt-4o-mini\",\n) -\u003e str:\n    response = client.chat.completions.create(\n        messages=[\n            {\n                \"role\": \"user\",\n                \"content\": [\n                    {\"type\": \"text\", \"text\": message},\n                    {\n                        \"type\": \"image_url\",\n                        \"image_url\": {\n                            \"url\": f\"data:image/jpeg;base64,{dogo_picture}\",\n                            \"detail\": \"low\",\n                        },\n                    },\n                ],\n            }\n        ],\n        model=model,\n    )\n    return response.choices[0].message.content\n\n# daidai takes care of loading dependencies \u0026 injecting assets!\nprint(ask(\"Hello, what's in the picture ?\"))\n# \u003e\u003e\u003e The picture features a dog with a black and white coat.\n\n# Or manage lifecycle with context manager for production usage\n# all predictors, assets and artifacts are automatically loaded and cleaned up\nwith ModelManager(preload=[ask]):\n    print(ask(\"Hello, what's in the picture ?\"))\n\n# or manually pass dependencies\nmy_other_openai_client = openai.OpenAI(timeout=0.1)\nprint(ask(\"Hello, what's in the picture ?\", client=my_other_openai_client))\n# \u003e\u003e\u003e openai.APITimeoutError: Request timed out.\n# OOOPS, the new client timed out, of course :-)\n```\n\nYou can visualize the dependency graph of the above code using the `daidai CLI`:\n\n```\n$ daidai list -m example.py\n\n📦 Daidai Components\n├── 📄 Artifacts\n│   └── https://images.pexels.com/photos/220938/pexels-photo-220938.jpeg\n│       ├── Cache strategies: no_cache\n│       └── Used by:\n│           └── dogo_picture (asset) as picture\n├── 🧩 Assets\n│   ├── openai_client\n│   └── dogo_picture\n│       └── Artifacts\n│           └── picture: https://images.pexels.com/photos/220938/pexels-photo-220938.jpeg - Cache: no_cache\n└── 🔮 Predictors\n    └── ask\n        └── Dependencies\n            ├── dogo_picture: dogo_picture (asset) - default\n            └── client: openai_client (asset) - timeout=5\n```\n\n## Roadmap\n\n- [x] Clean things up now that the UX has landed\n- [x] Protect file operations for parallelism / concurrency\n- [ ] Add docs\n- [ ] Add tests (unit, integration, e2e)\n- [ ] Add a cookbook with common patterns \u0026 recipes\n- [ ] Add support for async components\n- [ ] Enjoy the fruits of my labor 🍊\n\n# 🧠 Core Concepts\n\n`daidai` is built around a few key concepts that work together to provide a streamlined experience for developing and deploying ML components. The following explains these core concepts and how they interact.\n\nAt the heart of `daidai` are three types of components: Assets, Predictors and Artifacts.\n\n\u003e **TL;DR** Predictors are functions that perform computations, Assets are long-lived objects that are expensive to create and should be reused, and Artifacts are the raw files, model weights, and other resources that assets transform into usable components.\n\n## 🧩 Assets\n\nAssets represent long-lived objects that are typically expensive to create and should be reused across multiple operations, e.g.:\n_Loaded ML models (or parts of: weights etc.), Embedding models, Customer Configurations, Tokenizers, Database connections, API clients.._\n\nAssets have several important characteristics:\n\n1. They are computed once and cached, making them efficient for repeated use\n2. They can depend on other assets or artifacts\n3. They are automatically cleaned up when no longer needed\n4. They can implement resource cleanup through generator functions\n\nAssets are defined using the `@asset` decorator:\n\n```python\n@asset\ndef bert_model(\n    model_path: Annotated[Path, \"s3://models/bert-base.pt\"]\n) -\u003e BertModel:\n    return BertModel.from_pretrained(model_path)\n```\n\n## 🔮 Predictors\n\nPredictors are functions that use assets to perform actual computations or predictions. Unlike assets:\n\n1. They are not cached themselves\n2. They are meant to be called repeatedly with different inputs\n3. They can depend on multiple assets or even other predictors\n4. They focus on the business logic of your application\n\nPredictors are defined using the `@predictor` decorator:\n\n```python\n@predictor\ndef classify_text(\n    text: str,\n    model: Annotated[BertModel, bert_model],\n    tokenizer: Annotated[Tokenizer, tokenizer]\n) -\u003e str:\n    tokens = tokenizer(text)\n    prediction = model(tokens)\n    return prediction.label\n```\n\n\n## 📦 Artifacts\n\nArtifacts represent the fundamental building blocks in your ML system - the raw files, model weights, configuration data, and other persistent resources that assets transform into usable components.\nThey are discrete, versioned resources that can be stored, tracked, and managed across various storage systems.\n\nArtifacts have several important characteristics:\n\n1. They represent raw data resources that require minimal processing to retrieve\n2. They can be stored and accessed from a wide variety of storage systems\n3. They support flexible caching strategies to balance performance and resource usage\n4. They are automatically downloaded, cached, and managed by daidai\n5. They can be deserialized into various formats based on your needs\n\nArtifacts are defined using Python type annotations and are made vailable through the `artifacts` optional: `pip install daidai[artifacts]`\n\n```python\n\n@asset\ndef word_embeddings(\n    embeddings_file: Annotated[\n        Path,\n        \"s3://bucket/glove.txt\",  # Artifact location\n        {\"cache_strategy\": \"on_disk\"}\n    ]\n) -\u003e Dict[str, np.ndarray]:\n    with open(embeddings_file) as f:\n        embeddings = {...}\n        return embeddings\n```\n\n## 📦🧩🔮 Working Together\n\nThe true power of daidai emerges when `Artifacts`, `Assets`, and `Predictors` work together in a cohesive dependency hierarchy.\n\nThis architecture creates a natural progression from raw data to functional endpoints: `Artifacts` (raw files, model weights, configurations) are retrieved from storage systems, `Assets` transform these `artifacts` into reusable, long-lived objects, and `Predictors` utilize these `assets` to perform specific, **business-logic** computations.\n\nThis clean separation of concerns allows each component to focus on its specific role while forming part of a larger, integrated system.\n\n## 💉 Dependency Injection\n\n`daidai` uses a type-hint based dependency injection system that minimizes boilerplate while providing type safety. The system works as follows:\n\n### Type Annotations\n\nDependencies are declared using Python's `Annotated` type from the `typing` module:\n\n```python\nparam_name: Annotated[Type, Dependency, Optional[Configuration]]\n```\n\nWhere:\n\n- `Type` is the expected type of the parameter\n- `Dependency` is the function that will be called to obtain the dependency\n- `Optional[Configuration]` is an optional dictionary of configuration parameters\n\n### Automatic Resolution\n\nWhen you call a predictor or asset, `daidai` automatically:\n\n1. Identifies all dependencies (predictors, assets and artifacts)\n2. Resolves the dependency graph\n3. Loads or retrieves cached dependencies\n4. Injects them into your function\n\nThis happens transparently, so you can focus on your business logic rather than dependency management.\n\n\u003cdetails open\u003e\n\u003csummary\u003eSimple Dependency Resolution Flowchart\u003c/summary\u003e\n\nFor a single predictor with one asset dependency having one file dependency, the dependency resolution flow looks like this:\n\n\n```mermaid\nflowchart TD\n    A[User calls Predictor] --\u003e B{Predictor in cache?}\n\n    subgraph \"Dependency Resolution\"\n        B --\u003e|No| D[Resolve Dependencies]\n\n        subgraph \"Asset Resolution\"\n            D --\u003e E{Asset in cache?}\n            E --\u003e|No| G[Resolve Asset Dependencies]\n\n            subgraph \"File Handling\"\n                G --\u003e H{File in cache?}\n                H --\u003e|No| J[Download File]\n                J --\u003e K[Cache File based on Strategy]\n                K --\u003e I[Get Cached File]\n                H --\u003e|Yes| I\n            end\n\n            I --\u003e L[Deserialize File to Required Format]\n            L --\u003e M[Compute Asset with File]\n            M --\u003e N[Cache Asset]\n            N --\u003e F[Get Cached Asset]\n            E --\u003e|Yes| F\n        end\n\n        F --\u003e O[Create Predictor Partial Function]\n        O --\u003e P[Cache Prepared Predictor]\n    end\n\n    B --\u003e|Yes| C[Get Cached Predictor]\n    P --\u003e C\n\n    subgraph \"Execution\"\n        C --\u003e Q[Execute Predictor Function]\n    end\n\n    Q --\u003e R[Return Result to User]\n\n```\n\n\u003c/details\u003e\n\n### Manual Overrides\n\nYou can always override automatic dependency injection by explicitly passing values:\n\n```python\n# Normal automatic injection\nresult = classify_text(\"Sample text\")\n\n# Override with custom model\ncustom_model = load_my_custom_model()\nresult = classify_text(\"Sample text\", model=custom_model)\n```\n\nThis way, you can easily swap out components for testing, debugging, or A/B testing.\n\n## 📦 Artifacts (in-depth)\n\n### File Types\n\ndaidai supports various file types through type hints, allowing you to specify exactly how you want to interact with the artifact:\n\n- `Path`: Returns a Path object pointing to the downloaded file, ideal for when you need to work with the file using standard file operations\n- `str`: Returns the file content as a string, useful for text-based configurations or small text files\n- `bytes`: Returns the file content as bytes, perfect for binary data like images or serialized models\n- `TextIO`: Returns a text file handle (similar to `open(file, \"r\")`), best for streaming large text files\n- `BinaryIO`: Returns a binary file handle (similar to `open(file, \"rb\")`), ideal for streaming large binary files\n- `Generator[str]`: Returns a generator that yields lines from the file, optimal for processing large text files line by line\n- `Generator[bytes]`: Returns a generator that yields chunks of binary data, useful for processing large binary files in chunks\n\n### Cache Strategies\n\ndaidai offers multiple caching strategies for artifacts to balance performance, storage use, and reliability:\n\n- `on_disk`: Download once and keep permanently in the cache directory. Ideal for stable artifacts that change infrequently.\n- `on_disk_temporary`: Download to a temporary location, automatically deleted when the process exits. Best for large artifacts needed only for the current session.\n- `no_cache`: Do not cache the artifact, fetch it each time. Useful for dynamic content that changes frequently or when running in environments with limited write permissions.\n\n### Storage Systems\n\nThanks to `fsspec` integration, daidai supports a wide range of storage systems, allowing you to retrieve artifacts from virtually anywhere:\n\n- Local file system for development and testing\n- Amazon S3 for cloud-native workflows\n- Google Cloud Storage for GCP-based systems\n- Microsoft Azure Blob Storage for Azure environments\n- SFTP/FTP for legacy or on-premises data\n- HTTP/HTTPS for web-based resources\n- Hugging Face Hub for ML models and datasets\n- And many more through fsspec protocols\n\nThis unified access layer means your code remains the same regardless of where your artifacts are stored, making it easy to transition from local development to cloud deployment.\n\n## 🧹 Resource Lifecycle Management\n\n`daidai` automatically manages the lifecycle of resources to prevent leaks and ensure clean shutdown:\n\n### Automatic Cleanup\n\nFor basic resources, assets are automatically released when they're no longer needed. For resources requiring explicit cleanup (like database connections), `daidai` supports generator-based cleanup:\n\n```python\n@asset\ndef database_connection(db_url: str):\n    # Establish connection\n    conn = create_connection(db_url)\n    try:\n        yield conn  # Return the connection for use\n    finally:\n        conn.close()  # This runs during cleanup\n```\n\n### ModelManager\n\nThe `ModelManager` class provides explicit control over component lifecycle and is the recommended way for production usage:\n\n```python\n# Preload components and manage their lifecycle\nwith ModelManager(preload=[classify_text]) as manager:\n    # Components are ready to use\n    result = classify_text(\"Sample input\")\n    # More operations...\n# All resources are automatically cleaned up\n```\n\nModelManager features:\n\n- Preloading of components for predictable startup times\n- Namespace isolation for managing different environments\n- Explicit cleanup on exit\n- Support for custom configuration\n\n### Namespaces\n\n`daidai` supports isolating components into namespaces, which is useful for:\n\n- Running multiple model versions concurrently\n- Testing with different configurations\n- Implementing A/B testing\n\n```python\n# Production namespace\nwith ModelManager(preload=[model_v1], namespace=\"prod\"):\n    # Development namespace in the same process\n    with ModelManager(preload=[model_v2], namespace=\"dev\"):\n        # Both can be used without conflicts\n        prod_result = model_v1(\"input\")\n        dev_result = model_v2(\"input\")\n```\n\n### Caching and Performance\n\n`daidai` implements intelligent caching to optimize performance:\n\n- Assets are cached based on their configuration parameters\n- Artifacts use a content-addressed store for efficient storage\n- Memory usage is tracked (when pympler is installed, `pip install daidai[memory]`)\n- Cache invalidation is handled automatically based on dependency changes\n\nThis ensures your ML components load quickly while minimizing redundant computation and memory usage.\n\n## 🔧 Environment Configuration\n\n`daidai` can be configured through environment variables:\n\n- `DAIDAI_CACHE_DIR`: Directory for persistent file cache\n- `DAIDAI_CACHE_DIR_TMP`: Directory for temporary file cache\n- `DAIDAI_DEFAULT_CACHE_STRATEGY`: Default strategy for file caching, so you don't have to specify it for each file\n- `DAIDAI_FORCE_DOWNLOAD`: Force download even if cached versions exist\n- `DAIDAI_LOG_LEVEL`: Logging verbosity level\n\n## 🖥️ Command Line\n\ndaidai provides a CLI that helps you explore your components and their relationships, making it easier to understand your ML system's architecture \u0026 pipelines.\n\n```bash\n# Install the CLI\npip install daidai[cli]\n```\n\n### Commands\n\n#### List Components\n\nThe `list` command displays all daidai components in your module, showing assets, predictors, and artifacts along with their dependencies and configurations:\n\n```bash\n# List all components in a module\ndaidai list -m [module.py] [assets|predictors|artifacts] [-c cache_strategy] [-f format]\n\n```\nThe output provides a detailed view of your component graph, making it easy to visualize dependencies and configurations:\n\n```\n$ daidai list -m example.py\n\n📦 Daidai Components\n├── 📄 Artifacts\n│   └── s3://bucket/model.pt\n│       ├── Cache strategies: on_disk\n│       └── Used by:\n│           └── bert_model (asset) as model_path\n├── 🧩 Assets\n│   ├── bert_model\n│   │   └── Artifacts\n│   │       └── model_path: s3://bucket/model.pt - Cache: on_disk\n│   └── tokenizer\n│       └── Artifacts\n│           └── vocab_file: s3://bucket/vocab.txt - Cache: on_disk\n└── 🔮 Predictors\n    └── classify_text\n        └── Dependencies\n            ├── model: bert_model (asset) - default\n            └── tokenizer: tokenizer (asset) - default\n```\n\nThis visualization helps you understand:\n- Which artifacts are being used and their cache strategies\n- How assets depend on artifacts and other assets\n- How predictors compose multiple assets together\n\nThe CLI is particularly useful for:\n- Documenting your ML system architecture\n- Debugging dependency issues\n- Understanding resource usage patterns\n- Discovering optimization opportunities\n\nYou can also filter components by type or cache strategy, and customize the output format for easier integration with other tools.\n\n##### Scenario: pre-caching artifacts for faster startup\n\n```bash\ndaidai list -m example.py artifacts -c on_disk -f raw\n```\nThen, the **soon-to-be-built daidai cache command** will allow you to pre-cache / download all artifacts in the list to a specific directory, so you can build a Docker image with all dependencies pre-cached.\n\n```bash\n[previous command] | daidai cache -d /path/to/cache -\n```\n\nThen in your Dockerfile:\n\n```Dockerfile\nCOPY /path/to/cache /path/to/cache\nENV DAIDAI_CACHE_DIR /path/to/cache\n```\n\nEt voilà, your service will start up faster as all artifacts are already downloaded and cached.\n\n## 🧰 Adaptable Design\n\n`daidai` embraces an adaptable design philosophy that provides core functionality while allowing for extensive customization and extension. This approach enables you to integrate `daidai` into your existing ML infrastructure without forcing rigid patterns or workflows.\n\n### Pure Functions as Building Blocks\n\nAt its core, `daidai` uses pure functions decorated with `@asset` and `@predictor` rather than class hierarchies or complex abstractions:\n\n```python\n@asset\ndef embedding_model(model_path: Path) -\u003e Model:\n    return load_model(model_path)\n\n@predictor\ndef embed_text(text: str, model: Annotated[Model, embedding_model]) -\u003e np.ndarray:\n    return model.encode(text)\n```\n\nThis functional approach provides several advantages:\n\n1. **Composability**: Functions can be easily composed together to create complex pipelines\n2. **Testability**: Pure functions with explicit dependencies are straightforward to test\n3. **Transparency**: The data flow between components is clear and traceable\n4. **Interoperability**: Functions work with any Python object, not just specialized classes\n\n### Integration with External Systems\n\nYou can easily integrate `daidai` with external systems and frameworks:\n\n```python\n# Integration with existing ML experiment tracking\nimport mlflow\n@asset\ndef tracked_model(model_id: str, mlflow_uri: str) -\u003e Model:\n    mlflow.set_tracking_uri(mlflow_uri)\n    model_uri = f\"models:/{model_id}/Production\"\n    return mlflow.sklearn.load_model(model_uri)\n\n# Integration with metrics collection\n@predictor\ndef classified_with_metrics(\n    text: str,\n    model: Annotated[Model, classifier_model],\n    metrics_client: Annotated[MetricsClient, metrics]\n) -\u003e str:\n    result = model.predict(text)\n    metrics_client.increment(\"prediction_count\")\n    metrics_client.histogram(\"prediction_latency\", time.time() - start_time)\n    return result\n```\n\n### Adding Your Own Capabilities\n\n`daidai` can be extended with additional capabilities by composing with other libraries:\n\n### Input/Output Validation with Pydantic\n\n```python\n# Apply validation to predictor\n@predictor\n@validate_call(validate_return=True)\ndef analyze_sentiment(\n    text: TextInput,\n    model: Annotated[Model, sentiment_model],\n    min_length: int = 1\n) -\u003e SentimentResult:\n    # Input has been validated\n    result = model.predict(text)\n    return SentimentResult(\n        sentiment=result.label,\n        confidence=result.score,\n    ) # Output will be validated\n```\n\n### Performance Optimization with LRU Cache\n\n```python\nfrom functools import lru_cache\n\n@lru_cache(maxsize=1000)\n@predictor\ndef classify_text(\n    text: str,\n    model: Annotated[Model, classifier_model]\n) -\u003e str:\n    # This result will be cached based on text if only text is provided (and model injected)\n    return model.predict(text)\n\n```\n\n### Instrumentation and Observability\n\n```python\nfrom opentelemetry import trace\nimport time\nfrom functools import wraps\n\ntracer = trace.get_tracer(__name__)\n\ndef traced_predictor(func):\n    \"\"\"Decorator to add tracing to predictors\"\"\"\n    @wraps(func)\n    def wrapper(*args, **kwargs):\n        with tracer.start_as_current_span(func.__name__):\n            return func(*args, **kwargs)\n    return wrapper\n\ndef timed_predictor(func):\n    \"\"\"Decorator to measure and log execution time\"\"\"\n    @wraps(func)\n    def wrapper(*args, **kwargs):\n        start_time = time.perf_counter()\n        result = func(*args, **kwargs)\n        execution_time = time.perf_counter() - start_time\n        print(f\"{func.__name__} executed in {execution_time:.4f} seconds\")\n        return result\n    return wrapper\n\n@predictor\n@traced_predictor\n@timed_predictor\ndef predict_with_instrumentation(\n    text: str,\n    model: Annotated[Model, model]\n) -\u003e str:\n    # This call will be traced and timed\n    return model.predict(text)\n```\n\n### Replacing Components\n\nThe dependency injection system allows you to replace components at runtime without modifying any code:\n\n```python\n# Normal usage with automatic dependency resolution\nresult = embed_text(\"Example text\")\n\n# Replace the embedding model for A/B testing\nexperimental_model = load_experimental_model()\nresult_b = embed_text(\"Example text\", model=experimental_model)\n\n# Replace for a specific use case\nsmall_model = load_small_model()\nbatch_results = [embed_text(t, model=small_model) for t in large_batch]\n```\n\n`daidai`'s adaptable design ensures that you can build ML systems that meet your specific requirements while still benefiting from the core dependency management and caching features. Whether you're working on a simple prototype or a complex production system, `daidai` provides the flexibility to adapt to your needs without getting in your way.\n\n## 🧵 Concurrency \u0026 Parallelism\n\nWhile file operations are protected against race conditions (downloading, caching etc.), other operations **are not** due to the lazy nature of component loading.\nAs such, `daidai` cannot be considered thread-safe and does not plan to in the short term.\n\nHowever, there are ways to work around this limitation for multi-threaded applications:\n\n1. Create a shared `ModelManager` instance for all threads, but ensure that components are loaded before the threads are started:\n\n```python\n@asset\ndef model(model_path: Annotated[Path, \"s3://bucket/model.pkl\"]) -\u003e Model:\n    with open(model_path, \"rb\") as f:\n        return pickle.load(f)\n\n@predictor\ndef sentiment_classifier(text: str, model: Annotated[Model, model]):\n    return model.predict(text)\n\n\nwith ModelManager(preload=[sentiment_classifier]) as manager:\n    # sentiment_classifier and its dependencies (model) are loaded and\n    # ready to be used by all threads without issues\n    with ThreadPoolExecutor(max_workers=4) as executor:\n        results = list(executor.map(worker_function, data_chunks))\n```\n\n2. Create a separate `ModelManager` instance for each thread, each will benefit from the same disk cache but will not share components:\n\n```python\n# same predictor \u0026 asset definitions as above\n\ndef worker_function(data_chunk):\n    # Each thread has its own manager and namespace\n    with ModelManager(namespace=str(threading.get_ident())) as manager:\n        return my_predictor(data)\n\nwith ThreadPoolExecutor(max_workers=4) as executor:\n    results = list(executor.map(worker_function, data_chunks))\n```\n\nA few notes:\n\n- Creating separate ModelManager instances (approach #2) might lead to duplicate loading of the same components in memory across threads, while preloading (approach #1) ensures components are shared but requires knowing \u0026 loading all components in advance.\n- For most applications, approach #2 (separate managers) provides the safest experience, while approach #1 (preloading) is more memory-efficient and simple to implement for applications with large models.\n- Both approaches benefit from disk caching, so artifacts are only downloaded once regardless of how many ModelManager instances you create.\n\n## 🧪 Testing\n\n`daidai`'s design makes testing ML components straightforward and effective. The dependency injection pattern allows for clean separation of concerns and easy mocking of dependencies.\n\n### Unit Testing Components\n\nWhen unit testing assets or predictors, you can manually inject dependencies:\n\n```python\ndef test_text_classifier():\n    # Create a mock model that always returns \"positive\"\n    mock_model = lambda text: \"positive\"\n\n    # Pass the mock directly instead of using the real model\n    result = classify_text(\"Great product!\", model=mock_model)\n\n    assert result == \"positive\"\n```\n\n### Testing with Fixtures\n\nIn pytest, you can create fixtures that provide mock assets:\n\n```python\nimport pytest\n\n@pytest.fixture\ndef mock_embedding_model():\n    # Return a simplified embedding model for testing\n    return lambda text: np.ones(768) * 0.1\n\ndef test_semantic_search(mock_embedding_model):\n    # Use the fixture as a dependency\n    results = search_documents(\n        \"test query\",\n        embedding_model=mock_embedding_model\n    )\n    assert len(results) \u003e 0\n```\n\n### Integration Testing\n\nFor integration tests that verify the entire component pipeline:\n\n```python\n@pytest.fixture(scope=\"module\")\ndef test_model_manager():\n    # Set up a test namespace with real components\n    with ModelManager(\n        preload=[classify_text],\n        namespace=\"test\"\n    ) as manager:\n        yield manager\n\ndef test_end_to_end_classification(test_model_manager):\n    # This will use real components in the test namespace\n    result = classify_text(\"Test input\")\n    assert result in [\"positive\", \"negative\", \"neutral\"]\n```\n\n### Testing Artifacts\n\nFor artifacts, you can use local test files:\n\n```python\n@asset\ndef test_embeddings(\n    embeddings_file: Annotated[\n        Path,\n        \"file:///path/to/test_embeddings.npy\"\n    ]\n) -\u003e np.ndarray:\n    return np.load(embeddings_file)\n\n# In your test\ndef test_with_test_embeddings():\n    result = embed_text(\"test\", embeddings=test_embeddings())\n    assert result.shape == (768,)\n```\n\n`daidai`'s flexible design ensures that your ML components remain testable at all levels, from unit tests to integration tests, without requiring complex mocking frameworks or test setup.\n\n\n## 📚 Resources\n\n- [daidai 🍊 Documentation](https://antoinejeannot.github.io/daidai/)\n- [daidai 🍊 GitHub Repository](https://github.com/antoinejeannot/daidai/)\n- [daidai 🍊 PyPI Package](https://pypi.org/project/daidai/)\n\n\n## 📝 License\n\nThis project is licensed under the MIT License - see the [LICENSE](https://github.com/antoinejeannot/daidai/blob/main/LICENSE) file for details.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fantoinejeannot%2Fdaidai","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fantoinejeannot%2Fdaidai","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fantoinejeannot%2Fdaidai/lists"}