{"id":28324172,"url":"https://github.com/md-emon-hasan/informatruth","last_synced_at":"2026-07-25T06:31:35.564Z","repository":{"id":316370657,"uuid":"848113353","full_name":"Md-Emon-Hasan/InformaTruth","owner":"Md-Emon-Hasan","description":"Fine-tuned roberta-base classifier on the LIAR dataset. Aaccepts multiple input types text, URLs, and PDFs and outputs a prediction with a confidence score. It also leverages google/flan-t5-base to generate explanations and uses an Agentic AI with LangGraph to orchestrate agents for planning, retrieval, execution, fallback, and reasoning.","archived":false,"fork":false,"pushed_at":"2026-07-01T12:50:45.000Z","size":27342,"stargazers_count":1,"open_issues_count":0,"forks_count":1,"subscribers_count":1,"default_branch":"main","last_synced_at":"2026-07-01T13:25:51.382Z","etag":null,"topics":["ai-webapp","confidence-score","document-classification","end-to-end-ml-workflows","fake-news-detection","fine-tuning","flan-t5","huggingface-transformers","machine-learning","misinformation-detection","natural-language-processing","news-analysis","news-classification","roberta","sequence-classification","text-analysis","text-classification","transformers","truth-verification","url-parser"],"latest_commit_sha":null,"homepage":"https://informatruth.onrender.com","language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Md-Emon-Hasan.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2024-08-27T06:44:18.000Z","updated_at":"2026-07-01T12:50:50.000Z","dependencies_parsed_at":null,"dependency_job_id":"acc10feb-f8aa-4d2f-9cd8-3de1ac484dd1","html_url":"https://github.com/Md-Emon-Hasan/InformaTruth","commit_stats":null,"previous_names":["md-emon-hasan/informatruth"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/Md-Emon-Hasan/InformaTruth","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Md-Emon-Hasan%2FInformaTruth","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Md-Emon-Hasan%2FInformaTruth/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Md-Emon-Hasan%2FInformaTruth/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Md-Emon-Hasan%2FInformaTruth/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Md-Emon-Hasan","download_url":"https://codeload.github.com/Md-Emon-Hasan/InformaTruth/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Md-Emon-Hasan%2FInformaTruth/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":35869991,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-07-20T02:08:10.276Z","status":"online","status_checked_at":"2026-07-25T02:00:06.922Z","response_time":64,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["ai-webapp","confidence-score","document-classification","end-to-end-ml-workflows","fake-news-detection","fine-tuning","flan-t5","huggingface-transformers","machine-learning","misinformation-detection","natural-language-processing","news-analysis","news-classification","roberta","sequence-classification","text-analysis","text-classification","transformers","truth-verification","url-parser"],"created_at":"2025-05-25T17:10:31.964Z","updated_at":"2026-07-25T06:31:35.557Z","avatar_url":"https://github.com/Md-Emon-Hasan.png","language":"Jupyter Notebook","funding_links":[],"categories":[],"sub_categories":[],"readme":"# InformaTruth: Explainable AI Fake News Authenticity Analyzer\n[![CI/CD](https://github.com/Md-Emon-Hasan/InformaTruth/actions/workflows/main.yml/badge.svg)](https://github.com/Md-Emon-Hasan/InformaTruth/actions) [![Python](https://img.shields.io/badge/python-3.11-blue)](https://python.org) [![PyTorch](https://img.shields.io/badge/PyTorch-%23EE4C2C.svg?style=flat\u0026logo=PyTorch\u0026logoColor=white)](https://pytorch.org/) [![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Transformers-blue)](https://huggingface.co/) [![LangChain](https://img.shields.io/badge/LangChain-1C3C3C?style=flat\u0026logo=langchain\u0026logoColor=white)](https://python.langchain.com/) [![scikit-learn](https://img.shields.io/badge/scikit--learn-%23F7931E.svg?style=flat\u0026logo=scikit-learn\u0026logoColor=white)](https://scikit-learn.org/) [![Pandas](https://img.shields.io/badge/pandas-%23150458.svg?style=flat\u0026logo=pandas\u0026logoColor=white)](https://pandas.pydata.org/) [![FastAPI](https://img.shields.io/badge/FastAPI-005571?style=flat\u0026logo=fastapi)](https://fastapi.tiangolo.com) [![Docker](https://img.shields.io/badge/docker-%230db7ed.svg?style=flat\u0026logo=docker\u0026logoColor=white)](https://www.docker.com/) ![React](https://img.shields.io/badge/react-%2320232a.svg?style=flat\u0026logo=react\u0026logoColor=%2361DAFB) ![Tailwind CSS](https://img.shields.io/badge/tailwindcss-%2338B2AC.svg?style=flat\u0026logo=tailwind-css\u0026logoColor=white) ![Vite](https://img.shields.io/badge/vite-%23646CFF.svg?style=flat\u0026logo=vite\u0026logoColor=white)\n\nInformaTruth is an end-to-end AI-powered multi-agent fact-checking system that automatically verifies news articles, PDFs, and web content. It leverages QLoRA (4-bit) fine-tuning of RoBERTa, LangGraph orchestration, RAG pipelines, and fallback retrieval agents to deliver reliable, context-aware verification. The system features a modular multi-agent architecture including Planner, Retriever, Generator, Memory, and Fallback Agents, integrating diverse tools for comprehensive reasoning.\n\nIt achieves ~66% accuracy and ~62% macro-F1 (with ~78% recall on fake-news detection) on the LIAR dataset, with 95% query coverage and ~60% improved reliability through intelligent tool routing and memory integration. Designed for real-world deployment, InformaTruth includes a Flask-based responsive UI, FastAPI endpoints, Dockerized containers, and a CI/CD pipeline, enabling enterprise-grade automated fact verification at scale.\n\n[![Project demo video](https://github.com/user-attachments/assets/423ca9a1-caf1-405e-b671-be842d9a1240)](https://github.com/user-attachments/assets/423ca9a1-caf1-405e-b671-be842d9a1240)\n\n[![InformaTruth](https://github.com/user-attachments/assets/1e6717bc-53a3-4848-80a8-252c4eae8f5b)](https://github.com/user-attachments/assets/1e6717bc-53a3-4848-80a8-252c4eae8f5b)\n[![InformaTruth](https://github.com/user-attachments/assets/187a8cc1-75dc-46c5-809a-8d88214797e4)](https://github.com/user-attachments/assets/187a8cc1-75dc-46c5-809a-8d88214797e4)\n\n---\n\n## Live Demo\n\n**Try it now**: [InformaTruth — Fake News Detection AI App](https://informatruth.onrender.com)\n\n---\n\n## Tech Stack\n| **Category**                | **Technology/Resource**                                                                                |\n| --------------------------- | ------------------------------------------------------------------------------------------------------ |\n| **Core Framework**          | PyTorch, Transformers, HuggingFace                                                                     |\n| **Frontend Framework**      | **React.js** (Vite), Tailwind CSS, DaisyUI                                                             |\n| **Backend Framework**       | **FastAPI** (Async, Pydantic)                                                                          |\n| **Classification Model**    | QLoRA Fine-tuned RoBERTa-base (4-bit NF4 + LoRA adapters) on LIAR Dataset                              |\n| **Explanation Model**       | FLAN-T5-base (Zero-shot Prompting)                                                                     |\n| **Training Data**           | LIAR Dataset (Political Fact-Checking)                                                                 |\n| **Evaluation Metrics**      | Accuracy, Macro-F1, Per-class (Fake) Precision/Recall/F1, ROC-AUC                                      |\n| **Training Framework**      | HuggingFace Trainer                                                                                    |\n| **Fine-tuning Method**      | QLoRA — 4-bit NF4 quantized base + LoRA adapters (PEFT) on attention query/key/value projections       |\n| **LangGraph Orchestration** | LangGraph (Multi-Agent Directed Acyclic Execution Graph)                                               |\n| **Agents Used**             | PlannerAgent, InputHandlerAgent, ToolRouterAgent, ExecutorAgent, ExplanationAgent, FallbackSearchAgent |\n| **Input Modalities**        | Raw Text, Website URLs (via Newspaper3k), PDF Documents (via PyMuPDF)                                  |\n| **Tool Augmentation**       | DuckDuckGo Search API (Fallback), Wikipedia (Planned), ToolRouter Logic                                |\n| **Web Scraping**            | Newspaper3k (HTML → Clean Article)                                                                     |\n| **PDF Parsing**             | PyMuPDF (Backend) / PDF.js (Frontend)                                                                  |\n| **Explainability**          | Natural language justification generated using FLAN-T5                                                 |\n| **State Management**        | Shared State Object (LangGraph-compatible)                                                             |\n| **Hosting Platform**        | Render (Docker)                                                                                        |\n| **Version Control**         | Git, GitHub                                                                                            |\n| **Logging \u0026 Debugging**     | Centralized Logs in `logs/` directory                                                                  |\n| **Database**                | **SQLite** + **SQLModel** (Auto-persistence of analysis results)                                       |\n| **Input Support**           | Text, URLs, PDF documents                                                                              |\n\n---\n\n## Key Features\n\n* **Monolithic \u0026 Agentic Architecture**\n  Strictly organized codebase following agentic principles with modular separation of concerns.\n\n* **Modern React Frontend**\n  A responsive, pixel-perfect UI built with **React**, **Vite**, and **Tailwind CSS**, featuring dark mode and glassmorphism design.\n\n* **FastAPI Backend**\n  High-performance asynchronous API handling automatic documentation and efficient model serving.\n\n* **Multi-Format Input Support**\n  Accepts raw **text**, **web URLs**, and **PDF documents** (with client-side text extraction).\n\n* **Full NLP Pipeline**\n  Integrates **fake news classification** (QLoRA-tuned RoBERTa) and **natural language explanation** (FLAN-T5).\n\n* **Modular Agent-Based Architecture**\n  Built using **LangGraph** with modular agents: `Planner`, `Router`, `Executor`, and `Fallback`.\n\n* **Explanation Generation**\n  Uses **FLAN-T5** to generate human-readable rationales for model predictions.\n\n* **Comprehensive Testing**\n  Targeting 100% test coverage with automated unit and integration tests using `pytest` (Backend) and `vitest` (Frontend).\n\n* **Structured Logging**\n  All logs are automatically saved to the `logs/` directory for better debugging and monitoring.\n\n---\n\n## Project File Structure\n\n```bash\nInformaTruth/\n│\n├── .github/\n│   └── workflows/\n│       └── main.yml                  # CI/CD Configuration\n│\n├── backend/                          # FastAPI Backend\n│   ├── app/                          # Application Package\n│   │   ├── agents/                   # Modular Pipeline Agents\n│   │   │   ├── executor.py\n│   │   │   ├── fallback_search.py\n│   │   │   ├── input_handler.py\n│   │   │   ├── planner.py\n│   │   │   └── router.py\n│   │   ├── graph/                    # LangGraph Orchestration\n│   │   │   ├── builder.py\n│   │   │   └── state.py\n│   │   ├── models/                   # AI Model Wrappers\n│   │   │   ├── classifier.py\n│   │   │   ├── db.py                 # Database Models\n│   │   │   └── loader.py\n│   │   ├── utils/                    # Shared Utilities\n│   │   │   ├── logger.py\n│   │   │   └── results.py\n│   │   ├── db.py                     # Database Connection \u0026 Setup\n│   │   └── main.py                   # FastAPI entry point\n│   │   └── valid.tsv\n│   ├── logs/                         # Application Logs\n│   │   └── fake_news_pipeline.log\n│   ├── news/                         # Sample Data\n│   │   └── news.pdf\n│   ├── tests/                        # Backend Tests\n│   │   ├── conftest.py\n│   │   ├── test_agents.py\n│   │   ├── test_api.py\n│   │   ├── test_db.py                # Database Integration Tests\n│   │   ├── test_edge_cases.py\n│   │   ├── test_lifespan.py\n│   │   └── test_models.py\n│   ├── train/                        # Training module\n│   │   ├── config.py\n│   │   ├── data_loader.py\n│   │   ├── predictor.py\n│   │   ├── run.py\n│   │   ├── trainer.py\n│   │   └── utils.py\n│   ├── config.py                     # Global Configuration\n│   ├── Dockerfile                    # Backend Dockerfile\n│   ├── pyproject.toml                # Project Configuration\n│   ├── requirements.txt              # Python dependencies\n│   └── setup.py                      # Package Setup\n│\n├── frontend/                         # React Frontend\n│   ├── public/                       # Static Assets\n│   ├── src/                          # Source Code\n│   │   ├── assets/                   # Images/Vectors\n│   │   │   └── react.svg\n│   │   ├── components/               # React Components\n│   │   │   ├── AnalysisForm.jsx\n│   │   │   ├── Footer.jsx\n│   │   │   ├── Hero.jsx\n│   │   │   ├── Navbar.jsx\n│   │   │   └── ResultsParams.jsx\n│   │   ├── App.jsx                   # Main App Component\n│   │   ├── App.test.jsx              # Frontend Unit Tests\n│   │   ├── index.css                 # Tailwind\n│   │   ├── main.jsx                 \n│   │   └── setupTests.js             \n│   ├── Dockerfile                    # Frontend Dockerfile\n│   ├── eslint.config.js            \n│   ├── index.html                    # HTML Entry Point\n│   ├── package-lock.json           \n│   ├── package.json                 \n│   └── vite.config.js                # Vite Configuration\n│\n├── docker-compose.yml                # Docker Orchestration\n├── demo.mp4                          # Demo Video\n├── demo-1.png                        # Demo Image\n├── demo-2.png                        # Demo Image\n├── LICENSE                           # Project License\n├── README.md                         # Documentation\n├── render.yml                        # Render Deployment Config\n└── run.py                            # Root launch script\n```\n\n---\n\n## Getting Started\n\n### 1. Running the Application (Local Development)\nTo launch both the **FastAPI Backend** and **React Frontend** locally in parallel:\n```bash\npython run.py\n```\n- **Backend**: `http://localhost:8000`\n- **Frontend**: `http://localhost:5173`\n\n### 2. Running with Docker (Production/Containerized)\nTo build and run the entire stack using Docker Compose:\n```bash\ndocker-compose up --build\n```\n\n### 3. Running Training\nTo trigger the model training process (ensure you are in `backend/`):\n```bash\ncd backend\npython train/run.py\n```\n\n### 4. Running Tests\n**Backend (Pytest):**\n```bash\ncd backend\npython -m pytest tests/ --cov=app --cov-report=term-missing\n```\n\n**Frontend (Vitest):**\n```bash\ncd frontend\nnpm test\n```\n\n---\n\n## System Architecture\n```mermaid\ngraph TD\n    A[User Input] --\u003e B{Input Type}\n    B --\u003e|Text| C[Direct Text Processing]\n    B --\u003e|URL| D[Newspaper3k Parser]\n    B --\u003e|PDF| E[PDF.js Extraction]\n\n    C --\u003e F[Text Cleaner]\n    D --\u003e F\n    E --\u003e F\n\n    F --\u003e G[Context Validator]\n    G --\u003e|Sufficient Context| H[RoBERTa Classifier]\n    G --\u003e|Insufficient Context| I[Web Search Agent]\n    \n    I --\u003e J[Context Aggregator]\n    J --\u003e H\n\n    H --\u003e K[FLAN-T5 Explanation Generator]\n    K --\u003e L[Output Formatter]\n    \n    L --\u003e M[React Frontend]\n\n    style M fill:#e3f2fd,stroke:#90caf9\n    style G fill:#fff9c4,stroke:#fbc02d\n    style I fill:#fbe9e7,stroke:#ff8a65\n    style H fill:#f1f8e9,stroke:#aed581\n```\n\n---\n\n## Model Performance\n\nThe classifier is fine-tuned with **QLoRA** — a 4-bit NF4 quantized `roberta-base` with double quantization and bfloat16 compute, plus LoRA adapters (`r=16`, `alpha=32`, dropout `0.1`) on the attention query/key/value projections. Training runs for **5 epochs** (learning rate `2e-4`, batch size `16`, weight decay `0.01`). Only ~1.47M parameters (~1.17% of the model) are trainable; the quantized base stays frozen.\n\nPer-epoch validation metrics:\n\n| Epoch | Train Loss | Val Loss | Accuracy | Macro F1 | Precision (Fake) | Recall (Fake) | F1 (Fake) | ROC-AUC |\n|-------|------------|----------|----------|----------|------------------|---------------|-----------|---------|\n| 1     | 0.6395     | 0.6052   | 0.6807   | 0.6133   | 0.7374           | 0.8160        | 0.7747    | 0.6777  |\n| 2     | 0.6120     | 0.5835   | 0.6900   | 0.6049   | 0.7293           | 0.8576        | 0.7883    | 0.6837  |\n| 3     | 0.5907     | 0.5937   | 0.6659   | 0.6286   | 0.7630           | 0.7303        | 0.7463    | 0.7015  |\n| 4     | 0.5739     | 0.5847   | 0.6877   | 0.6284   | 0.7481           | 0.8079        | 0.7769    | 0.7017  |\n| 5     | 0.5553     | 0.5857   | 0.6877   | 0.6327   | 0.7525           | 0.7986        | 0.7748    | 0.7003  |\n\nFinal metrics on the held-out **LIAR test set**:\n\n| Metric              | Score  |\n|---------------------|--------|\n| Accuracy            | 0.6606 |\n| Macro F1            | 0.6162 |\n| Precision (Fake)    | 0.7205 |\n| Recall (Fake)       | 0.7751 |\n| F1 (Fake)           | 0.7468 |\n| ROC-AUC             | 0.6875 |\n| Eval Loss           | 0.6090 |\n\n\u003e Emphasis on **Recall** ensures the model catches most fake news cases.\n\n---\n\n## Professional Testing \u0026 Quality\n\n### 1. Linting\n**Backend (Ruff):**\n```bash\ncd backend\nruff check app/ tests/\n```\n**Frontend (ESLint):**\n```bash\ncd frontend\nnpm run lint\n```\n\n### 2. Code Formatting\n**Backend (Black):**\n```bash\ncd backend\nblack app/ tests/\n```\n\n---\n\n## CI/CD Pipeline (GitHub Actions)\nThe project utilizes a comprehensive **GitHub Actions** workflow for automated testing and validation.\n\n### Workflow Features:\n- **Backend**:\n  - Sets up Python 3.11\n  - Installs dependencies from `requirements.txt`\n  - Runs **Ruff** for linting\n  - Runs **Black** for formatting checks\n  - Executs **Pytest** with coverage reporting\n- **Frontend**:\n  - Sets up Node.js 18\n  - Installs dependencies (`npm ci`)\n  - Runs **ESLint**\n  - Executes **Vitest** unit tests\n- **Docker**:\n  - Builds `backend` and `frontend` Docker images upon successful tests\n\n### Trigger\nThe pipeline runs automatically on every `push` and `pull_request` to the `main` branch.\n\n---\n\n## **Developed By**\n\n**Md Emon Hasan**  \n**Email:** emon.mlengineer@gmail.com  \n**WhatsApp:** [+8801834363533](https://wa.me/8801834363533)  \n**Portfolio:** [Md-Emon-Hasan](https://emonlabs-ai.hitechparks.com/)  \n**GitHub:** [Md-Emon-Hasan](https://github.com/Md-Emon-Hasan)  \n**LinkedIn:** [Md Emon Hasan](https://www.linkedin.com/in/md-emon-hasan-695483237/)  \n**Facebook:** [Md Emon Hasan](https://www.facebook.com/mdemon.hasan2001/)\n\n---\n\n## License\nMIT License. Free to use with credit.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fmd-emon-hasan%2Finformatruth","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fmd-emon-hasan%2Finformatruth","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fmd-emon-hasan%2Finformatruth/lists"}