{"id":24251478,"url":"https://github.com/stevekgyang/mentalllama","last_synced_at":"2025-04-09T13:05:55.598Z","repository":{"id":196864382,"uuid":"695762947","full_name":"SteveKGYang/MentalLLaMA","owner":"SteveKGYang","description":"This repository introduces MentaLLaMA, the first open-source instruction following large language model for interpretable mental health analysis.","archived":false,"fork":false,"pushed_at":"2024-03-04T10:08:22.000Z","size":13885,"stargazers_count":251,"open_issues_count":4,"forks_count":27,"subscribers_count":6,"default_branch":"main","last_synced_at":"2025-04-08T09:41:06.076Z","etag":null,"topics":["chatgpt","gpt4","interpretability","language-model","large-language-models","llama2","mental-health","natural-language-processing","natural-language-understanding","social-media"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/SteveKGYang.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2023-09-24T06:25:41.000Z","updated_at":"2025-04-07T16:31:54.000Z","dependencies_parsed_at":null,"dependency_job_id":"8fb99d9b-0c3b-4f1d-8f01-54bd993a2471","html_url":"https://github.com/SteveKGYang/MentalLLaMA","commit_stats":{"total_commits":72,"total_committers":3,"mean_commits":24.0,"dds":"0.23611111111111116","last_synced_commit":"42cd29dfb91092843daae53ec162c75605d753ff"},"previous_names":["stevekgyang/mentalllama"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/SteveKGYang%2FMentalLLaMA","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/SteveKGYang%2FMentalLLaMA/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/SteveKGYang%2FMentalLLaMA/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/SteveKGYang%2FMentalLLaMA/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/SteveKGYang","download_url":"https://codeload.github.com/SteveKGYang/MentalLLaMA/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":248045231,"owners_count":21038553,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["chatgpt","gpt4","interpretability","language-model","large-language-models","llama2","mental-health","natural-language-processing","natural-language-understanding","social-media"],"created_at":"2025-01-15T02:50:56.935Z","updated_at":"2025-04-09T13:05:55.572Z","avatar_url":"https://github.com/SteveKGYang.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cp align=\"center\" width=\"100%\"\u003e\n\u003cimg src=\"https://i.postimg.cc/0Nd8VxbL/logo.png\"  width=\"100%\" height=\"100%\"\u003e\n\u003c/p\u003e\n\n\u003cdiv\u003e\n\u003cdiv align=\"left\"\u003e\n    \u003ca href='https://stevekgyang.github.io/' target='_blank'\u003eKailai Yang\u003csup\u003e1,2\u003c/sup\u003e\u0026emsp;\n    \u003ca href='https://www.zhangtianlin.top/' target='_blank'\u003eTianlin Zhang\u003csup\u003e1,2\u003c/sup\u003e\u0026emsp;\n    \u003ca target='_blank'\u003eShaoxiong Ji\u003csup\u003e3\u003c/sup\u003e\u003c/a\u003e\u0026emsp;\n    \u003ca target='_blank'\u003eQianqian Xie\u003csup\u003e1,2\u003c/sup\u003e\u003c/a\u003e\u0026emsp;\n    \u003ca target='_blank'\u003eZiyan Kuang\u003csup\u003e6\u003c/sup\u003e\u003c/a\u003e\u0026emsp;\n    \u003ca href='https://research.manchester.ac.uk/en/persons/sophia.ananiadou' target='_blank'\u003eSophia Ananiadou\u003csup\u003e1,2,4\u003c/sup\u003e\u003c/a\u003e\u0026emsp;\n    \u003ca target='_blank'\u003eJimin Huang\u003csup\u003e5\u003c/sup\u003e\u003c/a\u003e\n\u003c/div\u003e\n\u003cdiv\u003e\n\u003cdiv align=\"left\"\u003e\n    \u003csup\u003e1\u003c/sup\u003eNational Centre for Text Mining\u0026emsp;\n    \u003csup\u003e2\u003c/sup\u003eThe University of Manchester\u0026emsp;\n    \u003csup\u003e3\u003c/sup\u003eUniversity of Helsinki\u0026emsp;\n    \u003csup\u003e4\u003c/sup\u003eArtificial Intelligence Research Center, AIST\u0026emsp;\n    \u003csup\u003e5\u003c/sup\u003eWuhan University\u0026emsp;\n    \u003csup\u003e6\u003c/sup\u003eJiangxi Normal University\u0026emsp;\n\u003c/div\u003e\n\n\u003cdiv align=\"left\"\u003e\n    \u003cimg src='https://i.postimg.cc/Kj7RzvNr/nactem-hires.png' alt='NaCTeM' height='85px'\u003e\u0026emsp;\n    \u003cimg src='https://i.postimg.cc/nc2Jy6FN/uom.png' alt='UoM University Logo' height='85px'\u003e\u0026emsp;\n    \u003cimg src='https://i.postimg.cc/cJD3HsRY/helsinki.jpg' alt='helsinki Logo' height='85px'\u003e\u0026emsp;\n    \u003cimg src='https://i.postimg.cc/SNpxVKwg/airc-logo.png' alt='airc Logo' height='85px'\u003e\u0026emsp;\n    \u003cimg src='https://i.postimg.cc/CLtkBwz7/57-EDDD9-FB0-DF712-F3-AB627163-C2-1-EF15655-13-FCA.png' alt='Wuhan University Logo' height='85px'\u003e\n\u003c/div\u003e\n\n![](https://black.readthedocs.io/en/stable/_static/license.svg)\n\n## News\n📢 *Mar. 2, 2024* Full release of the test data for the IMHI benchmark.\n\n📢 *Feb. 1, 2024* Our MentaLLaMA paper: \n\"MentaLLaMA: Interpretable Mental Health Analysis on Social Media with Large Language Models\" has been accepted \nby WWW 2024!\n\n📢 *Oct. 31, 2023* We release the MentaLLaMA-33B-lora model, a 33B edition of MentaLLaMA based on \nVicuna-33B and the full IMHI dataset, but trained with LoRA due to the computational resources!\n\n📢 *Oct. 13, 2023* We release the training data for the following datasets: DR, dreaddit, SAD, \nMultiWD, and IRF. More to come, stay tuned!\n\n📢 *Oct. 7, 2023* Our evaluation paper: \n\"Towards Interpretable Mental Health Analysis with Large Language Models\" has been accepted \nby EMNLP 2023 main conference as a long paper!\n\n## Ethical Considerations\n\nThis repository and its contents are provided for **non-clinical research only**\n. None of the material constitutes actual diagnosis or advice, and help-seeker should get assistance\nfrom professional psychiatrists or clinical practitioners. No warranties, express or implied, are offered regarding the accuracy\n, completeness, or utility of the predictions and explanations. The authors and contributors are not\nresponsible for any errors, omissions, or any consequences arising from the use \nof the information herein. Users should exercise their own judgment and consult\nprofessionals before making any clinical-related decisions. The use\nof the software and information contained in this repository is entirely at the \nuser's own risk.\n\nThe raw datasets collected to build our IMHI dataset are from public\nsocial media platforms such as Reddit and Twitter, and we strictly\nfollow the privacy protocols and ethical principles to protect\nuser privacy and guarantee that anonymity is properly applied in\nall the mental health-related texts. In addition, to minimize misuse,\nall examples provided in our paper are paraphrased and obfuscated\nutilizing the moderate disguising scheme.\n\nIn addition, recent studies have indicated LLMs may introduce some potential\nbias, such as gender gaps. Meanwhile, some incorrect prediction results, inappropriate explanations, and over-generalization\nalso illustrate the potential risks of current LLMs. Therefore, there\nare still many challenges in applying the model to real-scenario\nmental health monitoring systems.\n\n*By using or accessing the information in this repository, you agree to indemnify, defend, and hold harmless the authors, contributors, and any affiliated organizations or persons from any and all claims or damages.*\n\n## Introduction\n\nThis project presents our efforts towards interpretable mental health analysis\nwith large language models (LLMs). In early works we comprehensively evaluate the zero-shot/few-shot \nperformances of the latest LLMs such as ChatGPT and GPT-4 on generating explanations\nfor mental health analysis. Based on the findings, we build the Interpretable Mental Health Instruction (IMHI)\ndataset with 105K instruction samples, the first multi-task and multi-source instruction-tuning dataset for interpretable mental\nhealth analysis on social media. Based on the IMHI dataset, We propose MentaLLaMA, the first open-source instruction-following LLMs for interpretable mental\nhealth analysis. MentaLLaMA can perform mental health\nanalysis on social media data and generate high-quality explanations for its predictions.\nWe also introduce the first holistic evaluation benchmark for interpretable mental health analysis with 19K test samples,\nwhich covers 8 tasks and 10 test sets. Our contributions are presented in these 2 papers:\n\n[The MentaLLaMA Paper](https://arxiv.org/abs/2309.13567) | [The Evaluation Paper](https://arxiv.org/abs/2304.03347)\n\n## MentaLLaMA Model \n\nWe provide 5 model checkpoints evaluated in the MentaLLaMA paper:\n\n- [MentaLLaMA-33B-lora](https://huggingface.co/klyang/MentaLLaMA-33B-lora): This model is fine-tuned based on the Vicuna-33B \nfoundation model and the full IMHI instruction tuning data. The training\ndata covers 8 mental health analysis tasks. The model can follow instructions to make accurate mental health analysis\nand generate high-quality explanations for the predictions. Due to the limitation of computational resources,\nwe train the MentaLLaMA-33B model with the PeFT technique LoRA, which significantly reduced memory usage.\n\n- [MentaLLaMA-chat-13B](https://huggingface.co/klyang/MentaLLaMA-chat-13B): This model is fine-tuned based on the Meta \nLLaMA2-chat-13B foundation model and the full IMHI instruction tuning data. The training\ndata covers 8 mental health analysis tasks. The model can follow instructions to make accurate mental health analysis\nand generate high-quality explanations for the predictions. Due to the model size, the inference\nare relatively slow.\n- [MentaLLaMA-chat-7B](https://huggingface.co/klyang/MentaLLaMA-chat-7B)|\n[MentaLLaMA-chat-7B-hf](https://huggingface.co/klyang/MentaLLaMA-chat-7B-hf): This model is fine-tuned based on the Meta \nLLaMA2-chat-7B foundation model and the full IMHI instruction tuning data. The training\ndata covers 8 mental health analysis tasks. The model can follow instructions to make mental health analysis\nand generate explanations for the predictions.\n- [MentalBART](https://huggingface.co/Tianlin668/MentalBART): This model is fine-tuned based on the BART-large foundation model\nand the full IMHI-completion data. The training data covers 8 mental health analysis tasks. The model cannot\nfollow instructions, but can make mental health analysis and generate explanations in a completion-based manner.\nThe smaller size of this model allows faster inference and easier deployment.\n- [MentalT5](https://huggingface.co/Tianlin668/MentalT5): This model is fine-tuned based on the T5-large foundation model\nand the full IMHI-completion data. The model cannot\nfollow instructions, but can make mental health analysis and generate explanations in a completion-based manner.\nThe smaller size of this model allows faster inference and easier deployment.\n\nYou can use the MentaLLaMA models in your Python project with the Hugging Face Transformers library. \nHere is a simple example of how to load the fully fine-tuned model:\n\n```python\nfrom transformers import LlamaTokenizer, LlamaForCausalLM\ntokenizer = LlamaTokenizer.from_pretrained(MODEL_PATH)\nmodel = LlamaForCausalLM.from_pretrained(MODEL_PATH, device_map='auto')\n```\n\nIn this example, LlamaTokenizer is used to load the tokenizer, and LlamaForCausalLM is used to load the model. The `device_map='auto'` argument is used to automatically\nuse the GPU if it's available. `MODEL_PATH` denotes your model save path.\n\nAfter loading the models, you can generate a response. Here is an example:\n\n```python\nprompt = 'Consider this post: \"work, it has been a stressful week! hope it gets better.\" Question: What is the stress cause of this post?'\ninputs = tokenizer(prompt, return_tensors=\"pt\")\n\n# Generate\ngenerate_ids = model.generate(inputs.input_ids, max_length=2048)\ntokenizer.batch_decode(generate_ids, skip_special_tokens=True, clean_up_tokenization_spaces=False)[0]\n```\n\nOur running of these codes on MentaLLaMA-chat-13B gets the following response:\n\n```\nAnswer: This post shows the stress cause related to work. Reasoning: The post explicitly mentions work as being stressful and expresses a hope that it gets better. This indicates that the poster is experiencing stress in relation to their work, suggesting that work is the primary cause of their stress in this instance.\n```\n\nFor the MentaLLaMA-33B-lora model, since our model is based on the Vicuna-33B foundation model, you need to first download the Vicuna-33B model [here](https://huggingface.co/lmsys/vicuna-33b-v1.3), \nand put it under the `./vicuna-33B` dir. Then download the MentaLLaMA-33B-lora weights and put it under the `./MentaLLaMA-33B-lora` dir.\n\n```python\nfrom peft import AutoPeftModelForCausalLM\nfrom transformers import AutoTokenizer\npeft_model = AutoPeftModelForCausalLM.from_pretrained(\"./MentaLLaMA-33B-lora\")\ntokenizer = AutoTokenizer.from_pretrained('./MentaLLaMA-33B-lora')\n```\n\nAfter loading the models, you can generate a response. Here is an example:\n\n```python\ninput_data = [\"Consider this post: I'm finally dead inside and I don't know how to feel about it Fear, Anger, Sadness... It's all gone. I just feel numb. Question: Does the poster suffer from depression?\"]\n\ninputs = tokenizer(input_data, return_tensors=\"pt\", padding=True)\ninput_ids = inputs.input_ids\n\ngenerate_ids = peft_model.generate(**inputs, max_length=2048)\n\ntruc_ids = generate_ids[0][len(input_ids[0]) :]\nresponse = tokenizer.decode(truc_ids, skip_special_tokens=True, spaces_between_special_tokens=False)\n```\n\nOur running of these codes on MentaLLaMA-33B-lora gets the following response:\n```\nReasoning: Yes, the poster suffers from depression. Reasoning: The poster's statement expresses a sense of emotional numbness and a lack of emotional response. This is a common symptom of depression, as individuals with depression often experience a diminished ability to feel emotions. The poster also mentions feeling dead inside, which further suggests a lack of emotional connection and a sense of hopelessness, both of which are common in depression. Overall, the language used and the description of emotional numbness align with symptoms commonly associated with depression.\n```\n\n## The IMHI Dataset\nWe collect raw data from 10 existing datasets covering 8 mental health analysis tasks, and transfer them into\ntest data for interpretable mental health analysis. Statistic about the 10 test sets are as follows:\n\n| Name                                                   | Task                                  | Data Split | Data Source | Annotation        | Released |\n|--------------------------------------------------------|---------------------------------------|------------|-------------|-------------------|----------|\n| [DR](https://aclanthology.org/W18-5903/)               | depression detection                  | 1,003/430/405        | Reddit      | Weak labels       | Yes      |\n| [CLP](https://aclanthology.org/W15-1204/)              | depression detection                  | 456/196/299        | Reddit      | Human annotations | Not yet  |\n| [dreaddit](https://aclanthology.org/D19-6213/)         | stress detection                      | 2,837/300/414        | Reddit      | Human annotations | Yes      |\n| [SWMH](https://arxiv.org/abs/2004.07601)               | mental disorders detection            | 34,822/8,705/10,882     | Reddit      | Weak labels       | Not yet  |\n| [T-SID](https://arxiv.org/abs/2004.07601)              | mental disorders detection            | 3,071/767/959        | Twitter     | Weak labels       | Not yet  |\n| [SAD](https://dl.acm.org/doi/10.1145/3411763.3451799)  | stress cause detection                | 5,547/616/684        | SMS         | Human annotations | Yes      |\n| [CAMS](https://aclanthology.org/2022.lrec-1.686/)      | depression/suicide cause detection    | 2,207/320/625        | Reddit      | Human annotations | Not yet  |\n| loneliness                                             | loneliness detection                  | 2,463/527/531        | Reddit      | Human annotations | Not yet  |\n| [MultiWD](https://github.com/drmuskangarg/MultiWD)     | Wellness dimensions detection         |  15,744/1,500/2,441      | Reddit      | Human annotations | Yes      |\n| [IRF](https://aclanthology.org/2023.findings-acl.757/) | Interpersonal risks factors detection | 3,943/985/2,113      | Reddit      | Human annotations | Yes      |\n\n### Training data\nWe introduce IMHI, the first multi-task and multi-source instruction-tuning dataset for interpretable mental\nhealth analysis on social media.\nWe currently release the training and evaluation data from the following sets: DR, dreaddit, SAD, MultiWD, and IRF. The instruction\ndata is put under\n```\n/train_data/instruction_data\n```\nThe items are easy to follow: the `query` row denotes the question, and the `gpt-3.5-turbo` row \ndenotes our modified and evaluated predictions and explanations from ChatGPT. `gpt-3.5-turbo` is used as\nthe golden response for evaluation.\n\nTo facilitate training on models with no instruction following ability, we also release part of the test data for \nIMHI-completion. The data is put under\n```\n/train_data/complete_data\n```\nThe file layouts are the same with instruction tuning data.\n\n### Evaluation Benchmark\nWe introduce the first holistic evaluation benchmark for interpretable mental health analysis with 19K test samples\n. All test data have been released. The instruction\ndata is put under\n```\n/test_data/test_instruction\n```\nThe items are easy to follow: the `query` row denotes the question, and the `gpt-3.5-turbo` row \ndenotes our modified and evaluated predictions and explanations from ChatGPT. `gpt-3.5-turbo` is used as\nthe golden response for evaluation.\n\nTo facilitate test on models with no instruction following ability, we also release part of the test data for \nIMHI-completion. The data is put under\n```\n/test_data/test_complete\n```\nThe file layouts are the same with instruction tuning data. \n\n## Model Evaluation\n\n### Response Generation\nTo evaluate your trained model on the IMHI benchmark, first load your model and generate responses for all\ntest items. We use the Hugging Face Transformers library to load the model. For LLaMA-based models, you can\ngenerate the responses with the following commands:\n```\ncd src\npython IMHI.py --model_path MODEL_PATH --batch_size 8 --model_output_path OUTPUT_PATH --test_dataset IMHI --llama --cuda\n```\n`MODEL_PATH` and `OUTPUT_PATH` denote the model save path and the save path for generated responses. \nAll generated responses will be put under `../model_output`. Some generated examples are shown in\n```\n./examples/response_generation_examples\n```\nYou can also evaluate with the IMHI-completion\ntest set with the following commands:\n```\ncd src\npython IMHI.py --model_path MODEL_PATH --batch_size 8 --model_output_path OUTPUT_PATH --test_dataset IMHI-completion --llama --cuda\n```\nYou can also load models that are not based on LLaMA by removing the `--llama` argument.\nIn the generated examples, the `goldens` row denotes the reference explanations and the `generated_text`\nrow denotes the generated responses from your model.\n\n### Correctness Evaluation\nThe first evaluation metric for our IMHI benchmark is to evaluate the classification correctness of the model\ngenerations. If your model can generate very regular responses, a rule-based classifier can do well to assign\na label to each response. We provide a rule-based classifier in `IMHI.py` and you can use it during the response\ngeneration process by adding the argument: `--rule_calculate` to your command. The classifier requires\nthe following template:\n\n```\n[label] Reasoning: [explanation]\n```\n\nHowever, as most LLMs are trained to generate diverse responses, a rule-based label classifier is impractical.\nFor example, MentaLLaMA can have the following response for an SAD query:\n```\nThis post indicates that the poster's sister has tested positive for ovarian cancer and that the family is devastated. This suggests that the cause of stress in this situation is health issues, specifically the sister's diagnosis of ovarian cancer. The post does not mention any other potential stress causes, making health issues the most appropriate label in this case.\n```\nTo solve this problem, in our [MentaLLaMA paper](https://arxiv.org/abs/2309.13567) we train 10 neural \nnetwork classifiers based on [MentalBERT](https://arxiv.org/abs/2110.15621), one for each collected raw dataset. The classifiers are trained to\nassign a classification label given the explanation. We release these 10 classifiers to facilitate future\nevaluations on IMHI benchmark.\n\nAll trained models achieve over 95% accuracy on the IMHI test data. Before you assign the labels, make sure \nyou have transferred your output files in the format of `/exmaples/response_generation_examples` and named\nas `DATASET.csv`. Put all the output files you want to label under the same DATA_PATH dir. Then download \nthe corresponding classifier models from the following links:\n\nThe models download links: [CAMS](https://huggingface.co/Tianlin668/CAMS), [CLP](https://huggingface.co/Tianlin668/CLP), [DR](https://huggingface.co/Tianlin668/DR),\n[dreaddit](https://huggingface.co/Tianlin668/dreaddit), [Irf](https://huggingface.co/Tianlin668/Irf), [loneliness](https://huggingface.co/Tianlin668/loneliness), [MultiWD](https://huggingface.co/Tianlin668/MultiWD),\n[SAD](https://huggingface.co/Tianlin668/SAD), [swmh](https://huggingface.co/Tianlin668/swmh), [t-sid](https://huggingface.co/Tianlin668/t-sid)\n\nPut all downloaded models under a MODEL_PATH dir and name each model with its dataset. For example, the model\nfor DR dataset should be put under `/MODEL_PATH/DR`. Now you can obtain the labels using these models with the following commands:\n```\ncd src\npython label_inference.py --model_path MODEL_PATH --data_path DATA_PATH --data_output_path OUTPUT_PATH --cuda\n```\nwhere `MODEL_PATH`, `DATA_PATH` denote your specified model and data dirs, and `OUTPUT_PATH` denotes your\noutput path. After processing, the output files should have the format as the examples in `/examples/label_data_examples`.\nIf you hope to calculate the metrics such as weight-F1 score and accuracy, add the argument `--calculate` to\nthe above command.\n### Explanation Quality Evaluation\nThe second evaluation metric for the IMHI benchmark is to evaluate the quality of the generated explanations.\nThe results in our \n[evaluation paper](https://arxiv.org/abs/2304.03347) show that [BART-score](https://arxiv.org/abs/2106.11520) is\nmoderately correlated with human annotations in 4 human evaluation aspects, and outperforms other automatic evaluation metrics. Therefore,\nwe utilize BART-score to evaluate the quality of the generated explanations. Specifically, you should first\ngenerate responses using the `IMHI.py` script and obtain the response dir as in `examples/response_generation_examples`.\nFirstly, download the [BART-score](https://github.com/neulab/BARTScore) directory and put it under `/src`, then\ndownload the [BART-score checkpoint](https://drive.google.com/file/d/1_7JfF7KOInb7ZrxKHIigTMR4ChVET01m/view?usp=sharing).\nThen score your responses with BART-score using the following commands:\n```\ncd src\npython score.py --gen_dir_name DIR_NAME --score_method bart_score --cuda\n```\n`DIR_NAME` denotes the dir name of your geenrated responses and should be put under `../model_output`. \nWe also provide other scoring methods. You can change `--score_method` to 'GPT3_score', 'bert_score', 'bleu', 'rouge'\nto use these metrics. For [GPT-score](https://github.com/jinlanfu/GPTScore), you need to first download\nthe project and put it under `/src`.\n\n## Human Annotations\n\nWe release our human annotations on AI-generated explanations to facilitate future research on aligning automatic evaluation\ntools for interpretable mental health analysis. Based on these human evaluation results, we tested various existing\nautomatic evaluation metrics on correlation with human preferences. The results in our \n[evaluation paper](https://arxiv.org/abs/2304.03347) show that \nBART-score is moderately correlated with human annotations in all 4 aspects.\n\n### Quality Evaluation\nIn our [evaluation paper](https://arxiv.org/abs/2304.03347), we manually labeled a subset of the AIGC results for the DR dataset in 4 aspects:\nfluency, completeness, reliability, and overall. The annotations are released in this dir:\n```\n/human_evaluation/DR_annotation\n```\nwhere we labeled 163 ChatGPT-generated explanations for the depression detection dataset DR. The file `chatgpt_data.csv`\nincludes 121 explanations that correctly classified by ChatGPT. `chatgpt_false_data.csv`\nincludes 42 explanations that falsely classified by ChatGPT. We also include 121 explanations that correctly \nclassified by InstructionGPT-3 in `gpt3_data.csv`.\n\n### Expert-written Golden Explanations\nIn our [MentaLLaMA paper](https://arxiv.org/abs/2309.13567), we invited one domain expert major in quantitative psychology\nto write an explanation for 350 selected posts (35 posts for each raw dataset). The golden set is used to accurately\nevaluate the explanation-generation ability of LLMs in an automatic manner. To facilitate future research, we\nrelease the expert-written explanations for the following datasets: DR, dreaddit, SWMH, T-SID, SAD, CAMS, \nloneliness, MultiWD, and IRF (35 samples each). The data is released in this dir:\n```\n/human_evaluation/test_instruction_expert\n```\nThe expert-written explanations are processed to follow the same format as other test datasets to facilitate\nmodel evaluations. You can test your model on the expert-written golden explanations with similar commands\nas in response generation. For example, you can test LLaMA-based models as follows:\n```\ncd src\npython IMHI.py --model_path MODEL_PATH --batch_size 8 --model_output_path OUTPUT_PATH --test_dataset expert --llama --cuda\n```\n\n\n## Citation\nIf you use the human annotations or analysis in the evaluation paper, please cite:\n\n```\n@inproceedings{yang2023towards,\n  title={Towards interpretable mental health analysis with large language models},\n  author={Yang, Kailai and Ji, Shaoxiong and Zhang, Tianlin and Xie, Qianqian and Kuang, Ziyan and Ananiadou, Sophia},\n  booktitle={Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing},\n  pages={6056--6077},\n  year={2023}\n}\n```\n\nIf you use MentaLLaMA in your work, please cite:\n```\n@article{yang2023mentalllama,\n  title={MentalLLaMA: Interpretable Mental Health Analysis on Social Media with Large Language Models},\n  author={Yang, Kailai and Zhang, Tianlin and Kuang, Ziyan and Xie, Qianqian and Ananiadou, Sophia},\n  journal={arXiv preprint arXiv:2309.13567},\n  year={2023}\n}\n```\n\n## License\n\nMentaLLaMA is licensed under [MIT]. Please find more details in the [MIT](LICENSE) file.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fstevekgyang%2Fmentalllama","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fstevekgyang%2Fmentalllama","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fstevekgyang%2Fmentalllama/lists"}