{"id":28907640,"url":"https://github.com/arbaznazir/datalineagepy","last_synced_at":"2026-04-28T17:02:24.490Z","repository":{"id":299570433,"uuid":"1003435454","full_name":"Arbaznazir/DataLineagePy","owner":"Arbaznazir","description":"86% faster data lineage tracking for pandas DataFrames with zero infrastructure. Real-time monitoring, ML anomaly detection, and enterprise compliance features.","archived":false,"fork":false,"pushed_at":"2025-09-17T11:25:31.000Z","size":3647,"stargazers_count":5,"open_issues_count":1,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-01-29T16:27:30.708Z","etag":null,"topics":["anomaly-detection","data-eng","data-governance","data-lineage","data-quality","data-science","dataframes","enterprise","etl","lineage-tracing","machine-learning","pandas","python"],"latest_commit_sha":null,"homepage":"https://pypi.org/project/datalineagepy/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Arbaznazir.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":"SECURITY_IMPLEMENTATION.md","support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2025-06-17T06:34:24.000Z","updated_at":"2026-01-22T22:15:06.000Z","dependencies_parsed_at":"2025-06-17T08:19:17.020Z","dependency_job_id":"6db21ba3-c7d4-4ab9-a6bd-5d5d9820c84c","html_url":"https://github.com/Arbaznazir/DataLineagePy","commit_stats":null,"previous_names":["arbaznazir/datalineagepy"],"tags_count":11,"template":false,"template_full_name":null,"purl":"pkg:github/Arbaznazir/DataLineagePy","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Arbaznazir%2FDataLineagePy","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Arbaznazir%2FDataLineagePy/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Arbaznazir%2FDataLineagePy/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Arbaznazir%2FDataLineagePy/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Arbaznazir","download_url":"https://codeload.github.com/Arbaznazir/DataLineagePy/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Arbaznazir%2FDataLineagePy/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":32390067,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-04-28T14:34:11.604Z","status":"ssl_error","status_checked_at":"2026-04-28T14:32:37.009Z","response_time":56,"last_error":"SSL_read: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["anomaly-detection","data-eng","data-governance","data-lineage","data-quality","data-science","dataframes","enterprise","etl","lineage-tracing","machine-learning","pandas","python"],"created_at":"2025-06-21T16:04:46.514Z","updated_at":"2026-04-28T17:02:24.485Z","avatar_url":"https://github.com/Arbaznazir.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# 🚀 DataLineagePy 3.0\n\n**Enterprise-Grade Python Data Lineage Tracking**\n\n[![Python 3.8+](https://img.shields.io/badge/python-3.8+-blue.svg)](https://www.python.org/downloads/)\n[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)\n[![Production Ready](https://img.shields.io/badge/status-production%20ready-green.svg)](https://github.com/Arbaznazir/DataLineagePy)\n[![Performance Score](https://img.shields.io/badge/performance-92.1%2F100-brightgreen.svg)](https://github.com/Arbaznazir/DataLineagePy)\n[![Enterprise Grade](https://img.shields.io/badge/enterprise-grade%20ready-gold.svg)](https://github.com/Arbaznazir/DataLineagePy)\n\n---\n\n\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"banner.jpg\" width=\"100%\" alt=\"DataLineagePy Banner\"/\u003e\n  \u003ch2\u003eBeautiful, Powerful, and Effortless Data Lineage for Python\u003c/h2\u003e\n  \u003cp\u003eTrack, visualize, and govern your data pipelines with zero friction.\u003c/p\u003e\n\u003c/div\u003e\n\n---\n\n## 🌟 Why DataLineagePy?\n\n- **Automatic, column-level lineage tracking** for all pandas DataFrames\n- **Enterprise performance**: memory-optimized, scalable, and production-ready\n- **Stunning visualizations**: interactive dashboards, HTML, PNG, SVG, and more\n- **Plug-and-play connectors**: MySQL, PostgreSQL, SQLite, and custom sources\n- **Security \u0026 compliance**: RBAC, AES-256 encryption, audit trails\n- **Real-time collaboration**: WebSocket server/client for team workflows\n- **ML/AI pipeline tracking**: Full auditability for machine learning steps\n- **Cloud-native deployment**: Docker, Kubernetes, Helm, Terraform\n\n---\n\n## 📋 Table of Contents\n\n- [Quick Start](#quick-start)\n- [Installation](#installation)\n- [Core Features](#core-features)\n- [Usage Guide](#usage-guide)\n- [Database Connectors](#database-connectors)\n- [Visualization \u0026 Reporting](#visualization--reporting)\n- [Performance Monitoring](#performance-monitoring)\n- [Security \u0026 Compliance](#security--compliance)\n- [ML/AI Pipeline Tracking](#mlai-pipeline-tracking)\n- [Enterprise Deployment](#enterprise-deployment)\n- [Use Cases](#use-cases)\n- [Documentation](#documentation)\n- [Contributing](#contributing)\n- [License](#license)\n\n---\n\n## 🚀 Quick Start\n\n```bash\npip install datalineagepy\n```\n\n```python\nfrom datalineagepy import LineageTracker, LineageDataFrame\nimport pandas as pd\n\ndf = pd.DataFrame({'a': [1, 2, 3], 'b': [4, 5, 6]})\ntracker = LineageTracker(name=\"demo\")\nldf = LineageDataFrame(df, name=\"my_df\", tracker=tracker)\nldf2 = ldf.filter(ldf._df['a'] \u003e 1)\nldf3 = ldf2.assign(c=ldf2._df['a'] + ldf2._df['b'])\ntracker.visualize()  # Interactive HTML dashboard\ntracker.export_lineage(\"lineage.json\")\n```\n\n---\n\n## 💾 Installation\n\n- **PyPI**: `pip install datalineagepy`\n- **With visualization**: `pip install datalineagepy[viz]`\n- **All features**: `pip install datalineagepy[all]`\n- **Conda**: `conda install -c conda-forge datalineagepy` _(coming soon)_\n- **Docker**: `docker pull datalineagepy/datalineagepy:latest`\n\nSee [Installation Guide](docs/installation.md) for advanced and enterprise setup.\n\n---\n\n## 📚 Core Features\n\n- **Automatic lineage tracking** for pandas DataFrames\n- **Data validation**: completeness, uniqueness, range, custom rules\n- **Profiling \u0026 analytics**: quality scoring, missing data, correlations\n- **Visualization**: HTML, PNG, SVG, interactive dashboards\n- **Performance monitoring**: execution time, memory, alerts\n- **Security**: RBAC, AES-256 encryption, audit trail\n- **Custom connectors**: SDK for any data source\n- **Versioning**: save, diff, rollback lineage graphs\n- **Collaboration**: real-time editing/viewing\n- **ML/AI pipeline tracking**: AutoMLTracker for full auditability\n\n---\n\n## 🔧 Usage Guide\n\n### 1. Lineage Tracking\n\n```python\nfrom datalineagepy import LineageTracker, LineageDataFrame\nimport pandas as pd\ntracker = LineageTracker(name=\"my_pipeline\")\ndf = pd.DataFrame({'x': [1,2,3], 'y': [4,5,6]})\nldf = LineageDataFrame(df, name=\"input\", tracker=tracker)\nldf2 = ldf.assign(z=ldf._df['x'] + ldf._df['y'])\nprint(tracker.export_graph())\n```\n\n### 2. Data Validation\n\n```python\nfrom datalineagepy.core.validation import DataValidator\nvalidator = DataValidator(tracker)\nrules = {'completeness': {'threshold': 0.9}, 'uniqueness': {'columns': ['x']}}\nresults = validator.validate_dataframe(ldf, rules)\nprint(results)\n```\n\n### 3. Profiling \u0026 Analytics\n\n```python\nfrom datalineagepy.core.analytics import DataProfiler\nprofiler = DataProfiler(tracker)\nprofile = profiler.profile_dataset(ldf, include_correlations=True)\nprint(profile)\n```\n\n### 4. Visualization \u0026 Reporting\n\n```python\nfrom datalineagepy.visualization.graph_visualizer import GraphVisualizer\nvisualizer = GraphVisualizer(tracker)\nvisualizer.generate_html(\"lineage.html\")\nvisualizer.generate_png(\"lineage.png\")\n```\n\n### 5. Performance Monitoring\n\n```python\nfrom datalineagepy.core.performance import PerformanceMonitor\nmonitor = PerformanceMonitor(tracker)\nmonitor.start_monitoring()\n_ = ldf._df.sum()\nmonitor.stop_monitoring()\nprint(monitor.get_performance_summary())\n```\n\n### 6. Security \u0026 Compliance\n\n```python\nfrom datalineagepy.security.rbac import RBACManager\nrbac = RBACManager()\nrbac.add_role('admin', ['read', 'write'])\nrbac.add_user('alice', ['admin'])\nprint(rbac.check_access('alice', 'write'))\n\nfrom datalineagepy.security.encryption.data_encryption import EncryptionManager\nimport os\nos.environ['MASTER_ENCRYPTION_KEY'] = 'supersecretkey1234567890123456'\nenc_mgr = EncryptionManager()\nsecret = 'Sensitive Data'\nencrypted = enc_mgr.encrypt_sensitive_data(secret)\ndecrypted = enc_mgr.decrypt_sensitive_data(encrypted)\nprint(decrypted)\n```\n\n### 7. Database Connectors\n\n```python\nfrom datalineagepy.connectors.database.mysql_connector import MySQLConnector\nfrom datalineagepy.core import LineageTracker\ndb_config = {'host': 'localhost', 'user': 'root', 'password': 'password', 'database': 'test_db'}\ntracker = LineageTracker()\nconn = MySQLConnector(**db_config, lineage_tracker=tracker)\nconn.execute_query('SELECT * FROM test_table')\nconn.close()\n```\n\n### 8. ML/AI Pipeline Tracking\n\n```python\nfrom datalineagepy import AutoMLTracker\ntracker = AutoMLTracker(name='ml_pipeline')\ntracker.log_step('fit', model='LogisticRegression', params={'solver': 'lbfgs'})\ntracker.log_step('predict', model='LogisticRegression')\nprint(tracker.export_ai_ready_format())\n```\n\n---\n\n## 📊 Visualization \u0026 Reporting\n\n- **Interactive HTML dashboards**: `tracker.visualize()`\n- **Export formats**: JSON, DOT, PNG, SVG, Excel, CSV\n- **Custom visualizations**: Use `GraphVisualizer` for advanced needs\n\n---\n\n## 🗄️ Database Connectors\n\n- **MySQL, PostgreSQL, SQLite**: Full lineage tracking for every query\n- **Custom connectors**: Build your own with the SDK\n- See [Database Connectors Guide](docs/user-guide/database-connectors.md)\n\n---\n\n## ⚡ Performance Monitoring\n\n- **Track execution time, memory, and operation stats**\n- **Alerting**: Slack, Email, custom hooks\n- **Production monitoring**: Integrate with Prometheus, Grafana, etc.\n\n---\n\n## 🔒 Security \u0026 Compliance\n\n- **RBAC**: Role-based access control for users and actions\n- **AES-256 encryption**: At-rest and in-transit data protection\n- **Audit trail**: Full operation history for compliance\n\n---\n\n## 🤖 ML/AI Pipeline Tracking\n\n- **AutoMLTracker**: Log, audit, and export every ML pipeline step\n- **Explainability**: Export pipeline steps for downstream analysis\n\n---\n\n## ☁️ Enterprise Deployment\n\n- **Docker, Kubernetes, Helm, Terraform**: Cloud-native ready\n- **Production scripts**: See `deploy/` for examples\n\n---\n\n## 💡 Use Cases\n\n- **Data science**: Reproducibility, experiment tracking, Jupyter integration\n- **Enterprise ETL**: Production pipelines, data quality, compliance\n- **Data governance**: Impact analysis, documentation, audit trails\n- **ML/AI**: Pipeline explainability, model audit, feature tracking\n\n---\n\n## 📖 Documentation\n\n- [User Guide](docs/user-guide/)\n- [API Reference](docs/api/)\n- [Quick Start](docs/quickstart.md)\n- [Enterprise Guide](docs/advanced/production.md)\n- [FAQ](docs/faq.md)\n- [Examples](examples/)\n\n---\n\n## 🤝 Contributing\n\nWe welcome contributions! Please see [CONTRIBUTING.md](CONTRIBUTING.md) for guidelines.\n\n---\n\n## 📄 License\n\nMIT License. See [LICENSE](LICENSE) for details.\n\n---\n\n\u003cdiv align=\"center\"\u003e\n  \u003cb\u003eDataLineagePy 3.0 \u0026mdash; The new standard for Python data lineage\u003c/b\u003e\u003cbr/\u003e\n  \u003ci\u003eBeautiful. Powerful. Effortless.\u003c/i\u003e\n\u003c/div\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Farbaznazir%2Fdatalineagepy","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Farbaznazir%2Fdatalineagepy","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Farbaznazir%2Fdatalineagepy/lists"}