Projects in Awesome Lists by GAIR-NLP
A curated list of projects in awesome lists by GAIR-NLP .
https://github.com/GAIR-NLP/O1-Journey
O1 Replication Journey: A Strategic Progress Report – Part I
Last synced: 30 Oct 2025
https://github.com/gair-nlp/factool
FacTool: Factuality Detection in Generative AI
chatgpt fact-checking generative-ai large-language-models natural-language-processing python
Last synced: 15 May 2025
https://github.com/GAIR-NLP/factool
FacTool: Factuality Detection in Generative AI
chatgpt fact-checking generative-ai large-language-models natural-language-processing python
Last synced: 05 Apr 2025
https://github.com/gair-nlp/factool?tab=readme-ov-file
FacTool: Factuality Detection in Generative AI
chatgpt fact-checking generative-ai large-language-models natural-language-processing python
Last synced: 29 Mar 2025
https://github.com/gair-nlp/anole
Anole: An Open, Autoregressive and Native Multimodal Models for Interleaved Image-Text Generation
Last synced: 07 Apr 2025
https://github.com/GAIR-NLP/anole
Anole: An Open, Autoregressive and Native Multimodal Models for Interleaved Image-Text Generation
Last synced: 05 Apr 2025
https://github.com/gair-nlp/openresearcher
OpenResearcher, an advanced Scientific Research Assistant
Last synced: 28 Jan 2026
https://github.com/gair-nlp/mathpile
[NeurlPS D&B 2024] Generative AI for Math: MathPile
corpus language-model large-language-models math pre-training
Last synced: 16 May 2025
https://github.com/GAIR-NLP/MathPile
[NeurlPS D&B 2024] Generative AI for Math: MathPile
corpus language-model large-language-models math pre-training
Last synced: 22 Jul 2025
https://github.com/gair-nlp/deepresearcher
Scaling Deep Research via Reinforcement Learning in Real-world Environments.
Last synced: 14 Jun 2025
https://github.com/GAIR-NLP/auto-j
Generative Judge for Evaluating Alignment
Last synced: 09 May 2025
https://github.com/GAIR-NLP/auto-j?tab=readme-ov-file
Generative Judge for Evaluating Alignment
Last synced: 29 Mar 2025
https://github.com/gair-nlp/auto-j
Generative Judge for Evaluating Alignment
Last synced: 13 Apr 2025
https://github.com/gair-nlp/pc-agent
PC Agent: While You Sleep, AI Works - A Cognitive Journey into Digital World
Last synced: 07 Apr 2025
https://github.com/GAIR-NLP/DeepResearcher
Scaling Deep Research via Reinforcement Learning in Real-world Environments.
Last synced: 01 May 2025
https://github.com/gair-nlp/prox
Offical Repo for "Programming Every Example: Lifting Pre-training Data Quality Like Experts at Scale"
continual continual-pre-training data-centric-ai data-quality llama llm mistral neural-symbolic pre-training
Last synced: 05 Apr 2025
https://github.com/gair-nlp/realign
Reformatted Alignment
alignment generative-ai large-language-models llms natural-language-processing nlp
Last synced: 13 Apr 2025
https://github.com/gair-nlp/olympicarena
[NeurIPS 2024] OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI
Last synced: 07 Apr 2025
https://github.com/gair-nlp/pc-agent-e
Efficient Agent Training for Computer Use
Last synced: 14 Jun 2025
https://gair-nlp.github.io/OlympicArena/
This is the official repository of the paper "OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI"
Last synced: 27 Oct 2025
https://github.com/gair-nlp/entropy-abf
Official implementation for 'Extending LLMs’ Context Window with 100 Samples'
Last synced: 13 Apr 2025
https://github.com/gair-nlp/octothinker
Revisiting Mid-training in the Era of RL Scaling
llama llm mid-training post-training pre-training qwen reasoning rl verl
Last synced: 30 Jun 2025
https://github.com/gair-nlp/maye
Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation Scheme
Last synced: 14 Jun 2025
https://github.com/gair-nlp/reasoneval
[AAAI 2025 oral] Evaluating Mathematical Reasoning Beyond Accuracy
Last synced: 13 Apr 2025
https://gair-nlp.github.io/benbench/
Benchmarking Benchmark Leakage in Large Language Models
benchmarks dataset large-language-models leakage-detection
Last synced: 29 Mar 2025
https://github.com/gair-nlp/benbench
Benchmarking Benchmark Leakage in Large Language Models
benchmarks dataset large-language-models leakage-detection
Last synced: 30 Oct 2025
https://github.com/GAIR-NLP/AgencyBench
AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts
Last synced: 08 Sep 2026
https://github.com/gair-nlp/scaleeval
Scalable Meta-Evaluation of LLMs as Evaluators
evaluation-framework generative-ai llm nlp
Last synced: 23 Jun 2025
https://github.com/gair-nlp/mops
[ACL 2024] Code for "MoPS: Modular Story Premise Synthesis for Open-Ended Automatic Story Generation"
Last synced: 25 Jun 2025
https://github.com/gair-nlp/behonest
BeHonest: Benchmarking Honesty in Large Language Models
alignment benchmark evaluation honesty llm nlp
Last synced: 13 Apr 2025
https://gair-nlp.github.io/BeHonest/
BeHonest: Benchmarking Honesty in Large Language Models
alignment benchmark evaluation honesty llm nlp
Last synced: 27 Oct 2025
https://github.com/gair-nlp/cognition-engineering
Generative AI Act II: Test Time Scaling Drives Cognition Engineering
Last synced: 22 Aug 2025
https://github.com/gair-nlp/safety-j
Safety-J: Evaluating Safety with Critique
Last synced: 13 Apr 2025
https://github.com/gair-nlp/thinking-with-generated-images
Doodling our way to AGI ✏️ 🖼️ 🧠
Last synced: 14 Jun 2025
https://github.com/gair-nlp/asi-arch
AlphaGo Moment for Model Architecture Discovery
Last synced: 28 Jul 2025
https://github.com/gair-nlp/megascience
MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
llama llm post-training qwen reasoning science scientific-reasoning
Last synced: 28 Jul 2025
https://github.com/gair-nlp/academiclaw
AcademiClaw: When Students Set Challenges for AI Agents — a bilingual benchmark of 80 university student-sourced academic tasks.
academic ai benchmark evaluation llm openclaw
Last synced: 28 Jun 2026
https://github.com/gair-nlp/lm-open-science-evaluation
Reproducible and flexible LLM evaluations for scientific reasoning.
evaluation llm reasoning science scientific-reasoning
Last synced: 28 Jul 2025