{"id":19499421,"url":"https://github.com/leondgarse/keras_cv_attention_models","last_synced_at":"2025-04-08T02:42:05.875Z","repository":{"id":37265036,"uuid":"391777965","full_name":"leondgarse/keras_cv_attention_models","owner":"leondgarse","description":"Keras beit,caformer,CMT,CoAtNet,convnext,davit,dino,efficientdet,edgenext,efficientformer,efficientnet,eva,fasternet,fastervit,fastvit,flexivit,gcvit,ghostnet,gpvit,hornet,hiera,iformer,inceptionnext,lcnet,levit,maxvit,mobilevit,moganet,nat,nfnets,pvt,swin,tinynet,tinyvit,uniformer,volo,vanillanet,yolor,yolov7,yolov8,yolox,gpt2,llama2, alias kecam","archived":false,"fork":false,"pushed_at":"2024-07-12T05:25:23.000Z","size":4430,"stargazers_count":587,"open_issues_count":6,"forks_count":92,"subscribers_count":23,"default_branch":"main","last_synced_at":"2024-09-07T09:35:41.211Z","etag":null,"topics":["attention","clip","coco","ddpm","detection","imagenet","keras","model","recognition","segment-anything","stable-diffusion","tensorflow","tf","tf2","visualizing"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/leondgarse.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2021-08-02T00:59:55.000Z","updated_at":"2024-08-16T17:55:00.000Z","dependencies_parsed_at":"2024-02-28T14:29:40.118Z","dependency_job_id":"c5712cdb-5493-491c-ad8d-71a19d1e6e16","html_url":"https://github.com/leondgarse/keras_cv_attention_models","commit_stats":{"total_commits":1023,"total_committers":5,"mean_commits":204.6,"dds":"0.012707722385141729","last_synced_commit":"f084b2cdcd204b860d7bbab8817f24c7d7784f3b"},"previous_names":[],"tags_count":138,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/leondgarse%2Fkeras_cv_attention_models","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/leondgarse%2Fkeras_cv_attention_models/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/leondgarse%2Fkeras_cv_attention_models/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/leondgarse%2Fkeras_cv_attention_models/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/leondgarse","download_url":"https://codeload.github.com/leondgarse/keras_cv_attention_models/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":247767232,"owners_count":20992538,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["attention","clip","coco","ddpm","detection","imagenet","keras","model","recognition","segment-anything","stable-diffusion","tensorflow","tf","tf2","visualizing"],"created_at":"2024-11-10T22:04:07.331Z","updated_at":"2025-04-08T02:42:05.848Z","avatar_url":"https://github.com/leondgarse.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# ___Keras_cv_attention_models___\n***\n- **WARNING: currently NOT compatible with `keras 3.x`, if using `tensorflow\u003e=2.16.0`, needs to install `pip install tf-keras~=$(pip show tensorflow | awk -F ': ' '/Version/{print $2}')` manually. While importing, import this package ahead of Tensorflow, or set `export TF_USE_LEGACY_KERAS=1`.**\n- **It's not recommended downloading and loading model from h5 file directly, better building model and loading weights like `import kecam; mm = kecam.models.LCNet050()`.**\n- **coco_train_script.py for TF is still under testing...**\n\u003c!-- TOC depthFrom:1 depthTo:6 withLinks:1 updateOnSave:1 orderedList:0 --\u003e\n\n- [___\u003e\u003e\u003e\u003e Roadmap and todo list \u003c\u003c\u003c\u003c___](https://github.com/leondgarse/keras_cv_attention_models/wiki/Roadmap)\n- [General Usage](#general-usage)\n  - [Basic](#basic)\n  - [T4 Inference](#t4-inference)\n  - [Layers](#layers)\n  - [Model surgery](#model-surgery)\n  - [ImageNet training and evaluating](#imagenet-training-and-evaluating)\n  - [COCO training and evaluating](#coco-training-and-evaluating)\n  - [CLIP training and evaluating](#clip-training-and-evaluating)\n  - [Text training](#text-training)\n  - [DDPM training](#ddpm-training)\n  - [Visualizing](#visualizing)\n  - [TFLite Conversion](#tflite-conversion)\n  - [Using PyTorch as backend](#using-pytorch-as-backend)\n  - [Using keras core as backend](#using-keras-core-as-backend)\n- [Recognition Models](#recognition-models)\n  - [AotNet](#aotnet)\n  - [BEiT](#beit)\n  - [BEiTV2](#beitv2)\n  - [BotNet](#botnet)\n  - [CAFormer](#caformer)\n  - [CMT](#cmt)\n  - [CoaT](#coat)\n  - [CoAtNet](#coatnet)\n  - [ConvNeXt](#convnext)\n  - [ConvNeXtV2](#convnextv2)\n  - [CoTNet](#cotnet)\n  - [CSPNeXt](#cspnext)\n  - [DaViT](#davit)\n  - [DiNAT](#dinat)\n  - [DINOv2](#dinov2)\n  - [EdgeNeXt](#edgenext)\n  - [EfficientFormer](#efficientformer)\n  - [EfficientFormerV2](#efficientformerv2)\n  - [EfficientNet](#efficientnet)\n  - [EfficientNetEdgeTPU](#efficientnetedgetpu)\n  - [EfficientNetV2](#efficientnetv2)\n  - [EfficientViT_B](#efficientvit_b)\n  - [EfficientViT_M](#efficientvit_m)\n  - [EVA](#eva)\n  - [EVA02](#eva02)\n  - [FasterNet](#fasternet)\n  - [FasterViT](#fastervit)\n  - [FastViT](#fastvit)\n  - [FBNetV3](#fbnetv3)\n  - [FlexiViT](#flexivit)\n  - [GCViT](#gcvit)\n  - [GhostNet](#ghostnet)\n  - [GhostNetV2](#ghostnetv2)\n  - [GMLP](#gmlp)\n  - [GPViT](#gpvit)\n  - [HaloNet](#halonet)\n  - [Hiera](#hiera)\n  - [HorNet](#hornet)\n  - [IFormer](#iformer)\n  - [InceptionNeXt](#inceptionnext)\n  - [LCNet](#lcnet)\n  - [LeViT](#levit)\n  - [MaxViT](#maxvit)\n  - [MetaTransFormer](#metatransformer)\n  - [MLP mixer](#mlp-mixer)\n  - [MobileNetV3](#mobilenetv3)\n  - [MobileViT](#mobilevit)\n  - [MobileViT_V2](#mobilevit_v2)\n  - [MogaNet](#moganet)\n  - [NAT](#nat)\n  - [NFNets](#nfnets)\n  - [PVT_V2](#pvt_v2)\n  - [RegNetY](#regnety)\n  - [RegNetZ](#regnetz)\n  - [RepViT](#repvit)\n  - [ResMLP](#resmlp)\n  - [ResNeSt](#resnest)\n  - [ResNetD](#resnetd)\n  - [ResNetQ](#resnetq)\n  - [ResNeXt](#resnext)\n  - [SwinTransformerV2](#swintransformerv2)\n  - [TinyNet](#tinynet)\n  - [TinyViT](#tinyvit)\n  - [UniFormer](#uniformer)\n  - [VanillaNet](#vanillanet)\n  - [VOLO](#volo)\n  - [WaveMLP](#wavemlp)\n- [Detection Models](#detection-models)\n  - [EfficientDet](#efficientdet)\n  - [YOLO_NAS](#yolo_nas)\n  - [YOLOR](#yolor)\n  - [YOLOV7](#yolov7)\n  - [YOLOV8](#yolov8)\n  - [YOLOX](#yolox)\n- [Language Models](#language-models)\n  - [GPT2](#gpt2)\n  - [LLaMA2](#llama2)\n- [Stable Diffusion](#stable-diffusion)\n- [Segmentation Models](#segmentation-models)\n  - [YOLOV8 Segmentation](#yolov8-segmentation)\n  - [Segment Anything](#segment-anything)\n- [Licenses](#licenses)\n- [Citing](#citing)\n\n\u003c!-- /TOC --\u003e\n***\n\n# General Usage\n## Basic\n  - **Default import** will not specific these while using them in READMEs.\n    ```py\n    import os\n    import sys\n    import tensorflow as tf\n    import numpy as np\n    import pandas as pd\n    import matplotlib.pyplot as plt\n    from tensorflow import keras\n    ```\n  - Install as pip package. `kecam` is a short alias name of this package. **Note**: the pip package `kecam` doesn't set any backend requirement, make sure either Tensorflow or PyTorch installed before hand. For PyTorch backend usage, refer [Keras PyTorch Backend](keras_cv_attention_models/pytorch_backend).\n    ```sh\n    pip install -U kecam\n    # Or\n    pip install -U keras-cv-attention-models\n    # Or\n    pip install -U git+https://github.com/leondgarse/keras_cv_attention_models\n    ```\n    Refer to each sub directory for detail usage.\n  - **Basic model prediction**\n    ```py\n    from keras_cv_attention_models import volo\n    mm = volo.VOLO_d1(pretrained=\"imagenet\")\n\n    \"\"\" Run predict \"\"\"\n    import tensorflow as tf\n    from tensorflow import keras\n    from keras_cv_attention_models.test_images import cat\n    img = cat()\n    imm = keras.applications.imagenet_utils.preprocess_input(img, mode='torch')\n    pred = mm(tf.expand_dims(tf.image.resize(imm, mm.input_shape[1:3]), 0)).numpy()\n    pred = tf.nn.softmax(pred).numpy()  # If classifier activation is not softmax\n    print(keras.applications.imagenet_utils.decode_predictions(pred)[0])\n    # [('n02124075', 'Egyptian_cat', 0.99664897),\n    #  ('n02123045', 'tabby', 0.0007249644),\n    #  ('n02123159', 'tiger_cat', 0.00020345),\n    #  ('n02127052', 'lynx', 5.4973923e-05),\n    #  ('n02123597', 'Siamese_cat', 2.675306e-05)]\n    ```\n    Or just use model preset `preprocess_input` and `decode_predictions`\n    ```py\n    from keras_cv_attention_models import coatnet\n    mm = coatnet.CoAtNet0()\n\n    from keras_cv_attention_models.test_images import cat\n    preds = mm(mm.preprocess_input(cat()))\n    print(mm.decode_predictions(preds))\n    # [[('n02124075', 'Egyptian_cat', 0.9999875), ('n02123045', 'tabby', 5.194884e-06), ...]]\n    ```\n    The preset `preprocess_input` and `decode_predictions` also compatible with PyTorch backend.\n    ```py\n    os.environ['KECAM_BACKEND'] = 'torch'\n\n    from keras_cv_attention_models import caformer\n    mm = caformer.CAFormerS18()\n    # \u003e\u003e\u003e\u003e Using PyTorch backend\n    # \u003e\u003e\u003e\u003e Aligned input_shape: [3, 224, 224]\n    # \u003e\u003e\u003e\u003e Load pretrained from: ~/.keras/models/caformer_s18_224_imagenet.h5\n\n    from keras_cv_attention_models.test_images import cat\n    preds = mm(mm.preprocess_input(cat()))\n    print(preds.shape)\n    # torch.Size([1, 1000])\n    print(mm.decode_predictions(preds))\n    # [[('n02124075', 'Egyptian_cat', 0.8817097), ('n02123045', 'tabby', 0.009335292), ...]]\n    ```\n  - **`num_classes=0`** set for excluding model top `GlobalAveragePooling2D + Dense` layers.\n    ```py\n    from keras_cv_attention_models import resnest\n    mm = resnest.ResNest50(num_classes=0)\n    print(mm.output_shape)\n    # (None, 7, 7, 2048)\n    ```\n  - **`num_classes={custom output classes}`** others than `1000` or `0` will just skip loading the header Dense layer weights. As `model.load_weights(weight_file, by_name=True, skip_mismatch=True)` is used for loading weights.\n    ```py\n    from keras_cv_attention_models import swin_transformer_v2\n\n    mm = swin_transformer_v2.SwinTransformerV2Tiny_window8(num_classes=64)\n    # \u003e\u003e\u003e\u003e Load pretrained from: ~/.keras/models/swin_transformer_v2_tiny_window8_256_imagenet.h5\n    # WARNING:tensorflow:Skipping loading weights for layer #601 (named predictions) due to mismatch in shape for weight predictions/kernel:0. Weight expects shape (768, 64). Received saved weight with shape (768, 1000)\n    # WARNING:tensorflow:Skipping loading weights for layer #601 (named predictions) due to mismatch in shape for weight predictions/bias:0. Weight expects shape (64,). Received saved weight with shape (1000,)\n    ```\n  - **Reload own model weights by set `pretrained=\"xxx.h5\"`**. Better than calling `model.load_weights` directly, if reloading model with different `input_shape` and with weights shape not matching.\n    ```py\n    import os\n    from keras_cv_attention_models import coatnet\n    pretrained = os.path.expanduser('~/.keras/models/coatnet0_224_imagenet.h5')\n    mm = coatnet.CoAtNet1(input_shape=(384, 384, 3), pretrained=pretrained)  # No sense, just showing usage\n    ```\n  - **Alias name `kecam`** can be used instead of `keras_cv_attention_models`. It's `__init__.py` only with `from keras_cv_attention_models import *`.\n    ```py\n    import kecam\n    mm = kecam.yolor.YOLOR_CSP()\n    imm = kecam.test_images.dog_cat()\n    preds = mm(mm.preprocess_input(imm))\n    bboxs, lables, confidences = mm.decode_predictions(preds)[0]\n    kecam.coco.show_image_with_bboxes(imm, bboxs, lables, confidences)\n    ```\n  - **Calculate flops** method from [TF 2.0 Feature: Flops calculation #32809](https://github.com/tensorflow/tensorflow/issues/32809#issuecomment-849439287). For PyTorch backend, needs `thop` `pip install thop`.\n    ```py\n    from keras_cv_attention_models import coatnet, resnest, model_surgery\n\n    model_surgery.get_flops(coatnet.CoAtNet0())\n    # \u003e\u003e\u003e\u003e FLOPs: 4,221,908,559, GFLOPs: 4.2219G\n    model_surgery.get_flops(resnest.ResNest50())\n    # \u003e\u003e\u003e\u003e FLOPs: 5,378,399,992, GFLOPs: 5.3784G\n    ```\n  - **[Deprecated] `tensorflow_addons`** is not imported by default. While reloading model depending on `GroupNormalization` like `MobileViTV2` from `h5` directly, needs to import `tensorflow_addons` manually first.\n    ```py\n    import tensorflow_addons as tfa\n\n    model_path = os.path.expanduser('~/.keras/models/mobilevit_v2_050_256_imagenet.h5')\n    mm = keras.models.load_model(model_path)\n    ```\n  - **Export TF model to onnx**. Needs `tf2onnx` for TF, `pip install onnx tf2onnx onnxsim onnxruntime`. For using PyTorch backend, exporting onnx is supported by PyTorch.\n    ```py\n    from keras_cv_attention_models import volo, nat, model_surgery\n    mm = nat.DiNAT_Small(pretrained=True)\n    model_surgery.export_onnx(mm, fuse_conv_bn=True, batch_size=1, simplify=True)\n    # Exported simplified onnx: dinat_small.onnx\n\n    # Run test\n    from keras_cv_attention_models.imagenet import eval_func\n    aa = eval_func.ONNXModelInterf(mm.name + '.onnx')\n    inputs = np.random.uniform(size=[1, *mm.input_shape[1:]]).astype('float32')\n    print(f\"{np.allclose(aa(inputs), mm(inputs), atol=1e-5) = }\")\n    # np.allclose(aa(inputs), mm(inputs), atol=1e-5) = True\n    ```\n  - **Model summary** `model_summary.csv` contains gathered model info.\n    - `params` for model params count in `M`\n    - `flops` for FLOPs in `G`\n    - `input` for model input shape\n    - `acc_metrics` means `Imagenet Top1 Accuracy` for recognition models, `COCO val AP` for detection models\n    - `inference_qps` for `T4 inference query per second` with `batch_size=1 + trtexec`\n    - `extra` means if any extra training info.\n    ```py\n    from keras_cv_attention_models import plot_func\n    plot_series = [\n        \"efficientnetv2\", 'tinynet', 'lcnet', 'mobilenetv3', 'fasternet', 'fastervit', 'ghostnet',\n        'inceptionnext', 'efficientvit_b', 'mobilevit', 'convnextv2', 'efficientvit_m', 'hiera',\n    ]\n    plot_func.plot_model_summary(\n        plot_series, model_table=\"model_summary.csv\", log_scale_x=True, allow_extras=['mae_in1k_ft1k']\n    )\n    ```\n    ![model_summary](https://github.com/leondgarse/keras_cv_attention_models/assets/5744524/0677c0a1-afa7-4b36-ab7b-dd160f0c2550)\n  - **Code format** is using `line-length=160`:\n    ```sh\n    find ./* -name \"*.py\" | grep -v __init__ | xargs -I {} black -l 160 {}\n    ```\n## T4 Inference\n  - **T4 Inference** in the model tables are tested using `trtexec` on `Tesla T4` with `CUDA=12.0.1-1, Driver=525.60.13`. All models are exported as ONNX using PyTorch backend, using `batch_szie=1` only. **Note: this data is for reference only, and vary in different batch sizes or benchmark tools or platforms or implementations**.\n  - All results are tested using colab [trtexec.ipynb](https://colab.research.google.com/drive/1xLwfvbZNqadkdAZu9b0UzOrETLo657oc?usp=drive_link). Thus reproducible by any others.\n  ```py\n  os.environ[\"KECAM_BACKEND\"] = \"torch\"\n\n  from keras_cv_attention_models import convnext, test_images, imagenet\n  # \u003e\u003e\u003e\u003e Using PyTorch backend\n  mm = convnext.ConvNeXtTiny()\n  mm.export_onnx(simplify=True)\n  # Exported onnx: convnext_tiny.onnx\n  # Running onnxsim.simplify...\n  # Exported simplified onnx: convnext_tiny.onnx\n\n  # Onnx run test\n  tt = imagenet.eval_func.ONNXModelInterf('convnext_tiny.onnx')\n  print(mm.decode_predictions(tt(mm.preprocess_input(test_images.cat()))))\n  # [[('n02124075', 'Egyptian_cat', 0.880507), ('n02123045', 'tabby', 0.0047998047), ...]]\n\n  \"\"\" Run trtexec benchmark \"\"\"\n  !trtexec --onnx=convnext_tiny.onnx --fp16 --allowGPUFallback --useSpinWait --useCudaGraph\n  ```\n## Layers\n  - [attention_layers](keras_cv_attention_models/attention_layers) is `__init__.py` only, which imports core layers defined in model architectures. Like `RelativePositionalEmbedding` from `botnet`, `outlook_attention` from `volo`, and many other `Positional Embedding Layers` / `Attention Blocks`.\n  ```py\n  from keras_cv_attention_models import attention_layers\n  aa = attention_layers.RelativePositionalEmbedding()\n  print(f\"{aa(tf.ones([1, 4, 14, 16, 256])).shape = }\")\n  # aa(tf.ones([1, 4, 14, 16, 256])).shape = TensorShape([1, 4, 14, 16, 14, 16])\n  ```\n## Model surgery\n  - [model_surgery](keras_cv_attention_models/model_surgery) including functions used to change model parameters after built.\n  ```py\n  from keras_cv_attention_models import model_surgery\n  mm = keras.applications.ResNet50()  # Trainable params: 25,583,592\n\n  # Replace all ReLU with PReLU. Trainable params: 25,606,312\n  mm = model_surgery.replace_ReLU(mm, target_activation='PReLU')\n\n  # Fuse conv and batch_norm layers. Trainable params: 25,553,192\n  mm = model_surgery.convert_to_fused_conv_bn_model(mm)\n  ```\n## ImageNet training and evaluating\n  - [ImageNet](keras_cv_attention_models/imagenet) contains more detail usage and some comparing results.\n  - [Init Imagenet dataset using tensorflow_datasets #9](https://github.com/leondgarse/keras_cv_attention_models/discussions/9).\n  - For custom dataset, `custom_dataset_script.py` can be used creating a `json` format file, which can be used as `--data_name xxx.json` for training, detail usage can be found in [Custom recognition dataset](https://github.com/leondgarse/keras_cv_attention_models/discussions/52#discussion-3971513).\n  - Another method creating custom dataset is using `tfds.load`, refer [Writing custom datasets](https://www.tensorflow.org/datasets/add_dataset) and [Creating private tensorflow_datasets from tfds #48](https://github.com/leondgarse/keras_cv_attention_models/discussions/48) by @Medicmind.\n  - Running an AWS Sagemaker estimator job using `keras_cv_attention_models` can be found in [AWS Sagemaker script example](https://github.com/leondgarse/keras_cv_attention_models/discussions/107) by @Medicmind.\n  - `aotnet.AotNet50` default parameters set is a typical `ResNet50` architecture with `Conv2D use_bias=False` and `padding` like `PyTorch`.\n  - Default parameters for `train_script.py` is like `A3` configuration from [ResNet strikes back: An improved training procedure in timm](https://arxiv.org/pdf/2110.00476.pdf) with `batch_size=256, input_shape=(160, 160)`.\n    ```sh\n    # `antialias` is default enabled for resize, can be turned off be set `--disable_antialias`.\n    CUDA_VISIBLE_DEVICES='0' TF_XLA_FLAGS=\"--tf_xla_auto_jit=2\" python3 train_script.py --seed 0 -s aotnet50\n    ```\n    ```sh\n    # Evaluation using input_shape (224, 224).\n    # `antialias` usage should be same with training.\n    CUDA_VISIBLE_DEVICES='1' python3 eval_script.py -m aotnet50_epoch_103_val_acc_0.7674.h5 -i 224 --central_crop 0.95\n    # \u003e\u003e\u003e\u003e Accuracy top1: 0.78466 top5: 0.94088\n    ```\n    ![aotnet50_imagenet](https://user-images.githubusercontent.com/5744524/163795114-b2441e5d-94d5-4310-826a-958426f1343e.png)\n  - **Restore from break point** by setting `--restore_path` and `--initial_epoch`, and keep other parameters same. `restore_path` is higher priority than `model` and `additional_model_kwargs`, also restore `optimizer` and `loss`. `initial_epoch` is mainly for learning rate scheduler. If not sure where it stopped, check `checkpoints/{save_name}_hist.json`.\n    ```py\n    import json\n    with open(\"checkpoints/aotnet50_hist.json\", \"r\") as ff:\n        aa = json.load(ff)\n    len(aa['lr'])\n    # 41 ==\u003e 41 epochs are finished, initial_epoch is 41 then, restart from epoch 42\n    ```\n    ```sh\n    CUDA_VISIBLE_DEVICES='0' TF_XLA_FLAGS=\"--tf_xla_auto_jit=2\" python3 train_script.py --seed 0 -r checkpoints/aotnet50_latest.h5 -I 41\n    # \u003e\u003e\u003e\u003e Restore model from: checkpoints/aotnet50_latest.h5\n    # Epoch 42/105\n    ```\n  - **`eval_script.py`** is used for evaluating model accuracy. [EfficientNetV2 self tested imagenet accuracy #19](https://github.com/leondgarse/keras_cv_attention_models/discussions/19) just showing how different parameters affecting model accuracy.\n    ```sh\n    # evaluating pretrained builtin model\n    CUDA_VISIBLE_DEVICES='1' python3 eval_script.py -m regnet.RegNetZD8\n    # evaluating pretrained timm model\n    CUDA_VISIBLE_DEVICES='1' python3 eval_script.py -m timm.models.resmlp_12_224 --input_shape 224\n\n    # evaluating specific h5 model\n    CUDA_VISIBLE_DEVICES='1' python3 eval_script.py -m checkpoints/xxx.h5\n    # evaluating specific tflite model\n    CUDA_VISIBLE_DEVICES='1' python3 eval_script.py -m xxx.tflite\n    ```\n  - **Progressive training** refer to [PDF 2104.00298 EfficientNetV2: Smaller Models and Faster Training](https://arxiv.org/pdf/2104.00298.pdf). AotNet50 A3 progressive input shapes `96 128 160`:\n    ```sh\n    CUDA_VISIBLE_DEVICES='1' TF_XLA_FLAGS=\"--tf_xla_auto_jit=2\" python3 progressive_train_script.py \\\n    --progressive_epochs 33 66 -1 \\\n    --progressive_input_shapes 96 128 160 \\\n    --progressive_magnitudes 2 4 6 \\\n    -s aotnet50_progressive_3_lr_steps_100 --seed 0\n    ```\n    ![aotnet50_progressive_160](https://user-images.githubusercontent.com/5744524/151286851-221ff8eb-9fe9-4685-aa60-4a3ba98c654e.png)\n  - Transfer learning with `freeze_backbone` or `freeze_norm_layers`: [EfficientNetV2B0 transfer learning on cifar10 testing freezing backbone #55](https://github.com/leondgarse/keras_cv_attention_models/discussions/55).\n  - [Token label train test on CIFAR10 #57](https://github.com/leondgarse/keras_cv_attention_models/discussions/57). **Currently not working as well as expected**. `Token label` is implementation of [Github zihangJiang/TokenLabeling](https://github.com/zihangJiang/TokenLabeling), paper [PDF 2104.10858 All Tokens Matter: Token Labeling for Training Better Vision Transformers](https://arxiv.org/pdf/2104.10858.pdf).\n## COCO training and evaluating\n  - **Currently still under testing**.\n  - [COCO](keras_cv_attention_models/coco) contains more detail usage.\n  - `custom_dataset_script.py` can be used creating a `json` format file, which can be used as `--data_name xxx.json` for training, detail usage can be found in [Custom detection dataset](https://github.com/leondgarse/keras_cv_attention_models/discussions/52#discussioncomment-2460664).\n  - Default parameters for `coco_train_script.py` is `EfficientDetD0` with `input_shape=(256, 256, 3), batch_size=64, mosaic_mix_prob=0.5, freeze_backbone_epochs=32, total_epochs=105`. Technically, it's any `pyramid structure backbone` + `EfficientDet / YOLOX header / YOLOR header` + `anchor_free / yolor / efficientdet anchors` combination supported.\n  - Currently 4 types anchors supported, parameter **`anchors_mode`** controls which anchor to use, value in `[\"efficientdet\", \"anchor_free\", \"yolor\", \"yolov8\"]`. Default `None` for `det_header` presets.\n  - **NOTE: `YOLOV8` has a default `regression_len=64` for bbox output length. Typically it's `4` for other detection models, for yolov8 it's `reg_max=16 -\u003e regression_len = 16 * 4 == 64`.**\n\n    | anchors_mode | use_object_scores | num_anchors | anchor_scale | aspect_ratios | num_scales | grid_zero_start |\n    | ------------ | ----------------- | ----------- | ------------ | ------------- | ---------- | --------------- |\n    | efficientdet | False             | 9           | 4            | [1, 2, 0.5]   | 3          | False           |\n    | anchor_free  | True              | 1           | 1            | [1]           | 1          | True            |\n    | yolor        | True              | 3           | None         | presets       | None       | offset=0.5      |\n    | yolov8       | False             | 1           | 1            | [1]           | 1          | False           |\n\n    ```sh\n    # Default EfficientDetD0\n    CUDA_VISIBLE_DEVICES='0' python3 coco_train_script.py\n    # Default EfficientDetD0 using input_shape 512, optimizer adamw, freezing backbone 16 epochs, total 50 + 5 epochs\n    CUDA_VISIBLE_DEVICES='0' python3 coco_train_script.py -i 512 -p adamw --freeze_backbone_epochs 16 --lr_decay_steps 50\n\n    # EfficientNetV2B0 backbone + EfficientDetD0 detection header\n    CUDA_VISIBLE_DEVICES='0' python3 coco_train_script.py --backbone efficientnet.EfficientNetV2B0 --det_header efficientdet.EfficientDetD0\n    # ResNest50 backbone + EfficientDetD0 header using yolox like anchor_free anchors\n    CUDA_VISIBLE_DEVICES='0' python3 coco_train_script.py --backbone resnest.ResNest50 --anchors_mode anchor_free\n    # UniformerSmall32 backbone + EfficientDetD0 header using yolor anchors\n    CUDA_VISIBLE_DEVICES='0' python3 coco_train_script.py --backbone uniformer.UniformerSmall32 --anchors_mode yolor\n\n    # Typical YOLOXS with anchor_free anchors\n    CUDA_VISIBLE_DEVICES='0' python3 coco_train_script.py --det_header yolox.YOLOXS --freeze_backbone_epochs 0\n    # YOLOXS with efficientdet anchors\n    CUDA_VISIBLE_DEVICES='0' python3 coco_train_script.py --det_header yolox.YOLOXS --anchors_mode efficientdet --freeze_backbone_epochs 0\n    # CoAtNet0 backbone + YOLOX header with yolor anchors\n    CUDA_VISIBLE_DEVICES='0' python3 coco_train_script.py --backbone coatnet.CoAtNet0 --det_header yolox.YOLOX --anchors_mode yolor\n\n    # Typical YOLOR_P6 with yolor anchors\n    CUDA_VISIBLE_DEVICES='0' python3 coco_train_script.py --det_header yolor.YOLOR_P6 --freeze_backbone_epochs 0\n    # YOLOR_P6 with anchor_free anchors\n    CUDA_VISIBLE_DEVICES='0' python3 coco_train_script.py --det_header yolor.YOLOR_P6 --anchors_mode anchor_free  --freeze_backbone_epochs 0\n    # ConvNeXtTiny backbone + YOLOR header with efficientdet anchors\n    CUDA_VISIBLE_DEVICES='0' python3 coco_train_script.py --backbone convnext.ConvNeXtTiny --det_header yolor.YOLOR --anchors_mode yolor\n    ```\n    **Note: COCO training still under testing, may change parameters and default behaviors. Take the risk if would like help developing.**\n  - **`coco_eval_script.py`** is used for evaluating model AP / AR on COCO validation set. It has a dependency `pip install pycocotools` which is not in package requirements. More usage can be found in [COCO Evaluation](keras_cv_attention_models/coco#evaluation).\n    ```sh\n    # EfficientDetD0 using resize method bilinear w/o antialias\n    CUDA_VISIBLE_DEVICES='1' python3 coco_eval_script.py -m efficientdet.EfficientDetD0 --resize_method bilinear --disable_antialias\n    # \u003e\u003e\u003e\u003e [COCOEvalCallback] input_shape: (512, 512), pyramid_levels: [3, 7], anchors_mode: efficientdet\n\n    # YOLOX using BGR input format\n    CUDA_VISIBLE_DEVICES='1' python3 coco_eval_script.py -m yolox.YOLOXTiny --use_bgr_input --nms_method hard --nms_iou_or_sigma 0.65\n    # \u003e\u003e\u003e\u003e [COCOEvalCallback] input_shape: (416, 416), pyramid_levels: [3, 5], anchors_mode: anchor_free\n\n    # YOLOR / YOLOV7 using letterbox_pad and other tricks.\n    CUDA_VISIBLE_DEVICES='1' python3 coco_eval_script.py -m yolor.YOLOR_CSP --nms_method hard --nms_iou_or_sigma 0.65 \\\n    --nms_max_output_size 300 --nms_topk -1 --letterbox_pad 64 --input_shape 704\n    # \u003e\u003e\u003e\u003e [COCOEvalCallback] input_shape: (704, 704), pyramid_levels: [3, 5], anchors_mode: yolor\n\n    # Specify h5 model\n    CUDA_VISIBLE_DEVICES='1' python3 coco_eval_script.py -m checkpoints/yoloxtiny_yolor_anchor.h5\n    # \u003e\u003e\u003e\u003e [COCOEvalCallback] input_shape: (416, 416), pyramid_levels: [3, 5], anchors_mode: yolor\n    ```\n  - **[Experimental] Training using PyTorch backend**\n    ```py\n    import os, sys, torch\n    os.environ[\"KECAM_BACKEND\"] = \"torch\"\n\n    from keras_cv_attention_models.yolov8 import train, yolov8\n    from keras_cv_attention_models import efficientnet\n\n    global_device = torch.device(\"cuda:0\") if torch.cuda.is_available() and int(os.environ.get(\"CUDA_VISIBLE_DEVICES\", \"0\")) \u003e= 0 else torch.device(\"cpu\")\n    # model Trainable params: 7,023,904, GFLOPs: 8.1815G\n    bb = efficientnet.EfficientNetV2B0(input_shape=(3, 640, 640), num_classes=0)\n    model = yolov8.YOLOV8_N(backbone=bb, classifier_activation=None, pretrained=None).to(global_device)  # Note: classifier_activation=None\n    # model = yolov8.YOLOV8_N(input_shape=(3, None, None), classifier_activation=None, pretrained=None).to(global_device)\n    ema = train.train(model, dataset_path=\"coco.json\", initial_epoch=0)\n    ```\n    ![yolov8_training](https://user-images.githubusercontent.com/5744524/235142289-cb6a4da0-1ea7-4261-afdd-03a3c36278b8.png)\n## CLIP training and evaluating\n  - [CLIP](keras_cv_attention_models/clip) contains more detail usage.\n  - `custom_dataset_script.py` can be used creating a `tsv` / `json` format file, which can be used as `--data_name xxx.tsv` for training, detail usage can be found in [Custom caption dataset](https://github.com/leondgarse/keras_cv_attention_models/discussions/52#discussioncomment-6516154).\n  - **Train using `clip_train_script.py on COCO captions`** Default `--data_path` is a testing one `datasets/coco_dog_cat/captions.tsv`.\n    ```sh\n    CUDA_VISIBLE_DEVICES=1 TF_XLA_FLAGS=\"--tf_xla_auto_jit=2\" python clip_train_script.py -i 160 -b 128 \\\n    --text_model_pretrained None --data_path coco_captions.tsv\n    ```\n    **Train Using PyTorch backend by setting `KECAM_BACKEND='torch'`**\n    ```sh\n    KECAM_BACKEND='torch' CUDA_VISIBLE_DEVICES=1 python clip_train_script.py -i 160 -b 128 \\\n    --text_model_pretrained None --data_path coco_captions.tsv\n    ```\n    ![clip_torch_tf](https://github.com/leondgarse/keras_cv_attention_models/assets/5744524/4cbc22e4-907d-4735-81a0-41e0fc17ebc5)\n## Text training\n  - Currently it's only a simple one modified from [Github karpathy/nanoGPT](https://github.com/karpathy/nanoGPT).\n  - **Train using `text_train_script.py`** As dataset is randomly sampled, needs to specify `steps_per_epoch`\n    ```sh\n    CUDA_VISIBLE_DEVICES=1 TF_XLA_FLAGS=\"--tf_xla_auto_jit=2\" python text_train_script.py -m LLaMA2_15M \\\n    --steps_per_epoch 8000 --batch_size 8 --tokenizer SentencePieceTokenizer\n    ```\n    **Train Using PyTorch backend by setting `KECAM_BACKEND='torch'`**\n    ```sh\n    KECAM_BACKEND='torch' CUDA_VISIBLE_DEVICES=1 python text_train_script.py -m LLaMA2_15M \\\n    --steps_per_epoch 8000 --batch_size 8 --tokenizer SentencePieceTokenizer\n    ```\n    **Plotting**\n    ```py\n    from keras_cv_attention_models import plot_func\n    hists = ['checkpoints/text_llama2_15m_tensorflow_hist.json', 'checkpoints/text_llama2_15m_torch_hist.json']\n    plot_func.plot_hists(hists, addition_plots=['val_loss', 'lr'], skip_first=3)\n    ```\n    ![text_tf_torch](https://github.com/leondgarse/keras_cv_attention_models/assets/5744524/0fc3dd08-bb20-47be-9267-d9cc35fba4c0)\n## DDPM training\n  - [Stable Diffusion](keras_cv_attention_models/stable_diffusion) contains more detail usage.\n  - **Note: Works better with PyTorch backend, Tensorflow one seems overfitted if training logger like `--epochs 200`, and evaluation runs ~5 times slower. [???]**\n  - **Dataset** can be a directory containing images for basic DDPM training using images only, or a recognition json file created following [Custom recognition dataset](https://github.com/leondgarse/keras_cv_attention_models/discussions/52#discussion-3971513), which will train using labels as instruction.\n    ```sh\n    python custom_dataset_script.py --train_images cifar10/train/ --test_images cifar10/test/\n    # \u003e\u003e\u003e\u003e total_train_samples: 50000, total_test_samples: 10000, num_classes: 10\n    # \u003e\u003e\u003e\u003e Saved to: cifar10.json\n    ```\n  - **Train using `ddpm_train_script.py on cifar10 with labels`** Default `--data_path` is builtin `cifar10`.\n    ```py\n    # Set --eval_interval 50 as TF evaluation is rather slow [???]\n    TF_XLA_FLAGS=\"--tf_xla_auto_jit=2\" CUDA_VISIBLE_DEVICES=1 python ddpm_train_script.py --eval_interval 50\n    ```\n    **Train Using PyTorch backend by setting `KECAM_BACKEND='torch'`**\n    ```py\n    KECAM_BACKEND='torch' CUDA_VISIBLE_DEVICES=1 python ddpm_train_script.py\n    ```\n    ![ddpm_unet_test_E100](https://github.com/leondgarse/keras_cv_attention_models/assets/5744524/861f4004-4496-4aff-ae9c-706f4c04fef2)\n## Visualizing\n  - [Visualizing](keras_cv_attention_models/visualizing) is for visualizing convnet filters or attention map scores.\n  - **make_and_apply_gradcam_heatmap** is for Grad-CAM class activation visualization.\n    ```py\n    from keras_cv_attention_models import visualizing, test_images, resnest\n    mm = resnest.ResNest50()\n    img = test_images.dog()\n    superimposed_img, heatmap, preds = visualizing.make_and_apply_gradcam_heatmap(mm, img, layer_name=\"auto\")\n    ```\n    ![](https://user-images.githubusercontent.com/5744524/148199374-4944800e-a1fb-4df2-b9ba-43ce3dde88f2.png)\n  - **plot_attention_score_maps** is model attention score maps visualization.\n    ```py\n    from keras_cv_attention_models import visualizing, test_images, botnet\n    img = test_images.dog()\n    _ = visualizing.plot_attention_score_maps(botnet.BotNetSE33T(), img)\n    ```\n    ![](https://user-images.githubusercontent.com/5744524/147209511-f5194d73-9e4c-457e-a763-45a4025f452b.png)\n## TFLite Conversion\n  - Currently `TFLite` not supporting `tf.image.extract_patches` / `tf.transpose with len(perm) \u003e 4`. Some operations could be supported in latest or `tf-nightly` version, like previously not supported `gelu` / `Conv2D with groups\u003e1` are working now. May try if encountering issue.\n  - More discussion can be found [Converting a trained keras CV attention model to TFLite #17](https://github.com/leondgarse/keras_cv_attention_models/discussions/17). Some speed testing results can be found [How to speed up inference on a quantized model #44](https://github.com/leondgarse/keras_cv_attention_models/discussions/44#discussioncomment-2348910).\n  - Functions like `model_surgery.convert_groups_conv2d_2_split_conv2d` and `model_surgery.convert_gelu_to_approximate` are not needed using up-to-date TF version.\n  - Not supporting `VOLO` / `HaloNet` models converting, cause they need a longer `tf.transpose` `perm`.\n  - **model_surgery.convert_dense_to_conv** converts all `Dense` layer with 3D / 4D inputs to `Conv1D` / `Conv2D`, as currently TFLite xnnpack not supporting it.\n    ```py\n    from keras_cv_attention_models import beit, model_surgery, efficientformer, mobilevit\n\n    mm = efficientformer.EfficientFormerL1()\n    mm = model_surgery.convert_dense_to_conv(mm)  # Convert all Dense layers\n    converter = tf.lite.TFLiteConverter.from_keras_model(mm)\n    open(mm.name + \".tflite\", \"wb\").write(converter.convert())\n    ```\n    | Model             | Dense, use_xnnpack=false  | Conv, use_xnnpack=false   | Conv, use_xnnpack=true    |\n    | ----------------- | ------------------------- | ------------------------- | ------------------------- |\n    | MobileViT_S       | Inference (avg) 215371 us | Inference (avg) 163836 us | Inference (avg) 163817 us |\n    | EfficientFormerL1 | Inference (avg) 126829 us | Inference (avg) 107053 us | Inference (avg) 107132 us |\n  - **model_surgery.convert_extract_patches_to_conv** converts `tf.image.extract_patches` to a `Conv2D` version:\n    ```py\n    from keras_cv_attention_models import cotnet, model_surgery\n    from keras_cv_attention_models.imagenet import eval_func\n\n    mm = cotnet.CotNetSE50D()\n    mm = model_surgery.convert_groups_conv2d_2_split_conv2d(mm)\n    # mm = model_surgery.convert_gelu_to_approximate(mm)  # Not required if using up-to-date TFLite\n    mm = model_surgery.convert_extract_patches_to_conv(mm)\n    converter = tf.lite.TFLiteConverter.from_keras_model(mm)\n    open(mm.name + \".tflite\", \"wb\").write(converter.convert())\n    test_inputs = np.random.uniform(size=[1, *mm.input_shape[1:]])\n    print(np.allclose(mm(test_inputs), eval_func.TFLiteModelInterf(mm.name + '.tflite')(test_inputs), atol=1e-7))\n    # True\n    ```\n  - **model_surgery.prepare_for_tflite** is just a combination of above functions:\n    ```py\n    from keras_cv_attention_models import beit, model_surgery\n\n    mm = beit.BeitBasePatch16()\n    mm = model_surgery.prepare_for_tflite(mm)\n    converter = tf.lite.TFLiteConverter.from_keras_model(mm)\n    open(mm.name + \".tflite\", \"wb\").write(converter.convert())\n    ```\n  - **Detection models** including `efficinetdet` / `yolox` / `yolor`, model can be converted a TFLite format directly. If need [DecodePredictions](https://github.com/leondgarse/keras_cv_attention_models/blob/main/keras_cv_attention_models/coco/eval_func.py#L8) also included in TFLite model, need to set `use_static_output=True` for `DecodePredictions`, as TFLite requires a more static output shape. Model output shape will be fixed as `[batch, max_output_size, 6]`. The last dimension `6` means `[bbox_top, bbox_left, bbox_bottom, bbox_right, label_index, confidence]`, and those valid ones are where `confidence \u003e 0`.\n    ```py\n    \"\"\" Init model \"\"\"\n    from keras_cv_attention_models import efficientdet\n    model = efficientdet.EfficientDetD0(pretrained=\"coco\")\n\n    \"\"\" Create a model with DecodePredictions using `use_static_output=True` \"\"\"\n    model.decode_predictions.use_static_output = True\n    # parameters like score_threshold / iou_or_sigma can be set another value if needed.\n    nn = model.decode_predictions(model.outputs[0], score_threshold=0.5)\n    bb = keras.models.Model(model.inputs[0], nn)\n\n    \"\"\" Convert TFLite \"\"\"\n    converter = tf.lite.TFLiteConverter.from_keras_model(bb)\n    open(bb.name + \".tflite\", \"wb\").write(converter.convert())\n\n    \"\"\" Inference test \"\"\"\n    from keras_cv_attention_models.imagenet import eval_func\n    from keras_cv_attention_models import test_images\n\n    dd = eval_func.TFLiteModelInterf(bb.name + \".tflite\")\n    imm = test_images.cat()\n    inputs = tf.expand_dims(tf.image.resize(imm, dd.input_shape[1:-1]), 0)\n    inputs = keras.applications.imagenet_utils.preprocess_input(inputs, mode='torch')\n    preds = dd(inputs)[0]\n    print(f\"{preds.shape = }\")\n    # preds.shape = (100, 6)\n\n    pred = preds[preds[:, -1] \u003e 0]\n    bboxes, labels, confidences = pred[:, :4], pred[:, 4], pred[:, -1]\n    print(f\"{bboxes = }, {labels = }, {confidences = }\")\n    # bboxes = array([[0.22825494, 0.47238672, 0.816262  , 0.8700745 ]], dtype=float32),\n    # labels = array([16.], dtype=float32),\n    # confidences = array([0.8309707], dtype=float32)\n\n    \"\"\" Show result \"\"\"\n    from keras_cv_attention_models.coco import data\n    data.show_image_with_bboxes(imm, bboxes, labels, confidences, num_classes=90)\n    ```\n## Using PyTorch as backend\n  - **Experimental** [Keras PyTorch Backend](keras_cv_attention_models/pytorch_backend).\n  - **Set os environment `export KECAM_BACKEND='torch'` to enable this PyTorch backend.**\n  - Currently supports most recognition and detection models except hornet*gf / nfnets / volo. For detection models, using `torchvision.ops.nms` while running prediction.\n  - **Basic model build and prediction**.\n    - Will load same `h5` weights as TF one if available.\n    - Note: `input_shape` will auto fit image data format. Given `input_shape=(224, 224, 3)` or `input_shape=(3, 224, 224)`, will both set to `(3, 224, 224)` if `channels_first`.\n    - Note: model is default set to `eval` mode.\n    ```py\n    os.environ['KECAM_BACKEND'] = 'torch'\n    from keras_cv_attention_models import res_mlp\n    mm = res_mlp.ResMLP12()\n    # \u003e\u003e\u003e\u003e Load pretrained from: ~/.keras/models/resmlp12_imagenet.h5\n    print(f\"{mm.input_shape = }\")\n    # mm.input_shape = [None, 3, 224, 224]\n\n    import torch\n    print(f\"{isinstance(mm, torch.nn.Module) = }\")\n    # isinstance(mm, torch.nn.Module) = True\n\n    # Run prediction\n    from keras_cv_attention_models.test_images import cat\n    print(mm.decode_predictions(mm(mm.preprocess_input(cat())))[0])\n    # [('n02124075', 'Egyptian_cat', 0.9597896), ('n02123045', 'tabby', 0.012809471), ...]\n    ```\n  - **Export typical PyTorch onnx / pth**.\n    ```py\n    import torch\n    torch.onnx.export(mm, torch.randn(1, 3, *mm.input_shape[2:]), mm.name + \".onnx\")\n\n    # Or by export_onnx\n    mm.export_onnx()\n    # Exported onnx: resmlp12.onnx\n\n    mm.export_pth()\n    # Exported pth: resmlp12.pth\n    ```\n  - **Save weights as h5**. This `h5` can also be loaded in typical TF backend model. Currently it's only weights without model structure supported.\n    ```py\n    mm.save_weights(\"foo.h5\")\n    ```\n  - **Training with compile and fit** Note: loss function arguments should be `y_true, y_pred`, while typical torch loss functions using `y_pred, y_true`.\n    ```py\n    import torch\n    from keras_cv_attention_models.backend import models, layers\n    mm = models.Sequential([layers.Input([3, 32, 32]), layers.Conv2D(32, 3), layers.GlobalAveragePooling2D(), layers.Dense(10)])\n    if torch.cuda.is_available():\n        _ = mm.to(\"cuda\")\n    xx = torch.rand([64, *mm.input_shape[1:]])\n    yy = torch.functional.F.one_hot(torch.randint(0, mm.output_shape[-1], size=[64]), mm.output_shape[-1]).float()\n    loss = lambda y_true, y_pred: (y_true - y_pred.float()).abs().mean()\n    # Will check kwargs for calling `self.train_compile` or `torch.nn.Module.compile`\n    mm.compile(optimizer=\"AdamW\", loss=loss, metrics='acc', grad_accumulate=4)\n    mm.fit(xx, yy, epochs=2, batch_size=4)\n    ```\n## Using keras core as backend\n  - **[Experimental] Set os environment `export KECAM_BACKEND='keras_core'` to enable this `keras_core` backend. Not using `keras\u003e3.0`, as still not compiling with TensorFlow==2.15.0**\n  - `keras-core` has its own backends, supporting tensorflow / torch / jax, by editting `~/.keras/keras.json` `\"backend\"` value.\n  - Currently most recognition models except `HaloNet` / `BotNet` supported, also `GPT2` / `LLaMA2` supported.\n  - **Basic model build and prediction**.\n    ```py\n    !pip install sentencepiece  # required for llama2 tokenizer\n    os.environ['KECAM_BACKEND'] = 'keras_core'\n    os.environ['KERAS_BACKEND'] = 'jax'\n    import kecam\n    print(f\"{kecam.backend.backend() = }\")\n    # kecam.backend.backend() = 'jax'\n    mm = kecam.llama2.LLaMA2_42M()\n    # \u003e\u003e\u003e\u003e Load pretrained from: ~/.keras/models/llama2_42m_tiny_stories.h5\n    mm.run_prediction('As evening fell, a maiden stood at the edge of a wood. In her hands,')\n    # \u003e\u003e\u003e\u003e Load tokenizer from file: ~/.keras/datasets/llama_tokenizer.model\n    # \u003cs\u003e\n    # As evening fell, a maiden stood at the edge of a wood. In her hands, she held a beautiful diamond. Everyone was surprised to see it.\n    # \"What is it?\" one of the kids asked.\n    # \"It's a diamond,\" the maiden said.\n    # ...\n    ```\n***\n\n# Recognition Models\n## AotNet\n  - [Keras AotNet](keras_cv_attention_models/aotnet) is just a `ResNet` / `ResNetV2` like framework, that set parameters like `attn_types` and `se_ratio` and others, which is used to apply different types attention layer. Works like `byoanet` / `byobnet` from `timm`.\n  - Default parameters set is a typical `ResNet` architecture with `Conv2D use_bias=False` and `padding` like `PyTorch`.\n  ```py\n  from keras_cv_attention_models import aotnet\n  # Mixing se and outlook and halo and mhsa and cot_attention, 21M parameters.\n  # 50 is just a picked number that larger than the relative `num_block`.\n  attn_types = [None, \"outlook\", [\"bot\", \"halo\"] * 50, \"cot\"],\n  se_ratio = [0.25, 0, 0, 0],\n  model = aotnet.AotNet50V2(attn_types=attn_types, se_ratio=se_ratio, stem_type=\"deep\", strides=1)\n  model.summary()\n  ```\n## BEiT\n  - [Keras BEiT](keras_cv_attention_models/beit) includes models from [PDF 2106.08254 BEiT: BERT Pre-Training of Image Transformers](https://arxiv.org/pdf/2106.08254.pdf).\n\n  | Model                      | Params  | FLOPs   | Input | Top1 Acc | T4 Inference |\n  | -------------------------- | ------- | ------- | ----- | -------- | ------------ |\n  | [BeitBasePatch16, 21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/beit/beit_base_patch16_224_imagenet21k-ft1k.h5)  | 86.53M  | 17.61G  | 224   | 85.240   | 321.226 qps  |\n  | - [21k_ft1k, 384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/beit/beit_base_patch16_384_imagenet21k-ft1k.h5)            | 86.74M  | 55.70G  | 384   | 86.808   | 164.705 qps  |\n  | [BeitLargePatch16, 21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/beit/beit_large_patch16_224_imagenet21k-ft1k.h5) | 304.43M | 61.68G  | 224   | 87.476   | 105.998 qps  |\n  | - [21k_ft1k, 384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/beit/beit_large_patch16_384_imagenet21k-ft1k.h5)            | 305.00M | 191.65G | 384   | 88.382   | 45.7307 qps  |\n  | - [21k_ft1k, 512](https://github.com/leondgarse/keras_cv_attention_models/releases/download/beit/beit_large_patch16_512_imagenet21k-ft1k.h5)            | 305.67M | 363.46G | 512   | 88.584   | 21.3097 qps  |\n## BEiTV2\n  - [Keras BEiT](keras_cv_attention_models/beit) includes models from BeitV2 Paper [PDF 2208.06366 BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers](https://arxiv.org/pdf/2208.06366.pdf).\n\n  | Model              | Params  | FLOPs  | Input | Top1 Acc | T4 Inference |\n  | ------------------ | ------- | ------ | ----- | -------- | ------------ |\n  | BeitV2BasePatch16  | 86.53M  | 17.61G | 224   | 85.5     | 322.52 qps   |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/beit/beit_v2_base_patch16_224_imagenet21k-ft1k.h5) | 86.53M          | 17.61G | 224   | 86.5     | 322.52 qps   |\n  | BeitV2LargePatch16 | 304.43M | 61.68G | 224   | 87.3     | 105.734 qps  |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/beit/beit_v2_large_patch16_224_imagenet21k-ft1k.h5)         | 304.43M | 61.68G | 224   | 88.4     | 105.734 qps  |\n## BotNet\n  - [Keras BotNet](keras_cv_attention_models/botnet) is for [PDF 2101.11605 Bottleneck Transformers for Visual Recognition](https://arxiv.org/pdf/2101.11605.pdf).\n\n  | Model         | Params | FLOPs  | Input | Top1 Acc | T4 Inference |\n  | ------------- | ------ | ------ | ----- | -------- | ------------ |\n  | BotNet50      | 21M    | 5.42G  | 224   |          | 746.454 qps  |\n  | BotNet101     | 41M    | 9.13G  | 224   |          | 448.102 qps  |\n  | BotNet152     | 56M    | 12.84G | 224   |          | 316.671 qps  |\n  | [BotNet26T](https://github.com/leondgarse/keras_cv_attention_models/releases/download/botnet/botnet26t_256_imagenet.h5)     | 12.5M  | 3.30G  | 256   | 79.246   | 1188.84 qps  |\n  | [BotNextECA26T](https://github.com/leondgarse/keras_cv_attention_models/releases/download/botnet/botnext_eca26t_256_imagenet.h5) | 10.59M | 2.45G  | 256   | 79.270   | 1038.19 qps  |\n  | [BotNetSE33T](https://github.com/leondgarse/keras_cv_attention_models/releases/download/botnet/botnet_se33t_256_imagenet.h5)   | 13.7M  | 3.89G  | 256   | 81.2     | 610.429 qps  |\n## CAFormer\n  - [Keras CAFormer](keras_cv_attention_models/caformer) is for [PDF 2210.13452 MetaFormer Baselines for Vision](https://arxiv.org/pdf/2210.13452.pdf). `CAFormer` is using 2 transformer stacks, while `ConvFormer` is all conv blocks.\n\n  | Model                   | Params | FLOPs | Input | Top1 Acc | T4 Inference |\n  | ----------------------- | ------ | ----- | ----- | -------- | ------------ |\n  | [CAFormerS18](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/caformer_s18_224_imagenet.h5)             | 26M    | 4.1G  | 224   | 83.6     | 399.127 qps  |\n  | - [384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/caformer_s18_384_imagenet.h5)                   | 26M    | 13.4G | 384   | 85.0     | 181.993 qps  |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/caformer_s18_224_imagenet21k-ft1k.h5)      | 26M    | 4.1G  | 224   | 84.1     | 399.127 qps  |\n  | - [21k_ft1k, 384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/caformer_s18_384_imagenet21k-ft1k.h5) | 26M    | 13.4G | 384   | 85.4     | 181.993 qps  |\n  | [CAFormerS36](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/caformer_s36_224_imagenet.h5)             | 39M    | 8.0G  | 224   | 84.5     | 204.328 qps  |\n  | - [384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/caformer_s36_384_imagenet.h5)                   | 39M    | 26.0G | 384   | 85.7     | 102.04 qps   |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/caformer_s36_224_imagenet21k-ft1k.h5)      | 39M    | 8.0G  | 224   | 85.8     | 204.328 qps  |\n  | - [21k_ft1k, 384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/caformer_s36_384_imagenet21k-ft1k.h5) | 39M    | 26.0G | 384   | 86.9     | 102.04 qps   |\n  | [CAFormerM36](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/caformer_m36_224_imagenet.h5)             | 56M    | 13.2G | 224   | 85.2     | 162.257 qps  |\n  | - [384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/caformer_m36_384_imagenet.h5)                   | 56M    | 42.0G | 384   | 86.2     | 65.6188 qps  |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/caformer_m36_224_imagenet21k-ft1k.h5)      | 56M    | 13.2G | 224   | 86.6     | 162.257 qps  |\n  | - [21k_ft1k, 384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/caformer_m36_384_imagenet21k-ft1k.h5) | 56M    | 42.0G | 384   | 87.5     | 65.6188 qps  |\n  | [CAFormerB36](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/caformer_b36_224_imagenet.h5)             | 99M    | 23.2G | 224   | 85.5     | 116.865 qps  |\n  | - [384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/caformer_b36_384_imagenet.h5)                   | 99M    | 72.2G | 384   | 86.4     | 50.0244 qps  |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/caformer_b36_224_imagenet21k-ft1k.h5)      | 99M    | 23.2G | 224   | 87.4     | 116.865 qps  |\n  | - [21k_ft1k, 384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/caformer_b36_384_imagenet21k-ft1k.h5) | 99M    | 72.2G | 384   | 88.1     | 50.0244 qps  |\n\n  | Model                   | Params | FLOPs | Input | Top1 Acc | T4 Inference |\n  | ----------------------- | ------ | ----- | ----- | -------- | ------------ |\n  | [ConvFormerS18](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/convformer_s18_224_imagenet.h5)           | 27M    | 3.9G  | 224   | 83.0     | 295.114 qps  |\n  | - [384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/convformer_s18_384_imagenet.h5)                   | 27M    | 11.6G | 384   | 84.4     | 145.923 qps  |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/convformer_s18_224_imagenet21k-ft1k.h5)      | 27M    | 3.9G  | 224   | 83.7     | 295.114 qps  |\n  | - [21k_ft1k, 384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/convformer_s36_384_imagenet21k-ft1k.h5) | 27M    | 11.6G | 384   | 85.0     | 145.923 qps  |\n  | [ConvFormerS36](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/convformer_s36_224_imagenet.h5)           | 40M    | 7.6G  | 224   | 84.1     | 161.609 qps  |\n  | - [384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/convformer_s36_384_imagenet.h5)                   | 40M    | 22.4G | 384   | 85.4     | 80.2101 qps  |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/convformer_s36_224_imagenet21k-ft1k.h5)      | 40M    | 7.6G  | 224   | 85.4     | 161.609 qps  |\n  | - [21k_ft1k, 384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/convformer_s36_384_imagenet21k-ft1k.h5) | 40M    | 22.4G | 384   | 86.4     | 80.2101 qps  |\n  | [ConvFormerM36](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/convformer_m36_224_imagenet.h5)           | 57M    | 12.8G | 224   | 84.5     | 130.161 qps  |\n  | - [384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/convformer_m36_384_imagenet.h5)                   | 57M    | 37.7G | 384   | 85.6     | 63.9712 qps  |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/convformer_m36_224_imagenet21k-ft1k.h5)      | 57M    | 12.8G | 224   | 86.1     | 130.161 qps  |\n  | - [21k_ft1k, 384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/convformer_m36_384_imagenet21k-ft1k.h5) | 57M    | 37.7G | 384   | 86.9     | 63.9712 qps  |\n  | [ConvFormerB36](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/convformer_b36_224_imagenet.h5)           | 100M   | 22.6G | 224   | 84.8     | 98.0751 qps  |\n  | - [384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/convformer_b36_384_imagenet.h5)                   | 100M   | 66.5G | 384   | 85.7     | 48.5897 qps  |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/convformer_b36_224_imagenet21k-ft1k.h5)      | 100M   | 22.6G | 224   | 87.0     | 98.0751 qps  |\n  | - [21k_ft1k, 384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/caformer/convformer_b36_384_imagenet21k-ft1k.h5) | 100M   | 66.5G | 384   | 87.6     | 48.5897 qps  |\n## CMT\n  - [Keras CMT](keras_cv_attention_models/cmt) is for [PDF 2107.06263 CMT: Convolutional Neural Networks Meet Vision Transformers](https://arxiv.org/pdf/2107.06263.pdf).\n\n  | Model                              | Params | FLOPs | Input | Top1 Acc | T4 Inference |\n  | ---------------------------------- | ------ | ----- | ----- | -------- | ------------ |\n  | CMTTiny, (Self trained 105 epochs) | 9.5M   | 0.65G | 160   | 77.4     | 315.566 qps  |\n  | - [(305 epochs)](https://github.com/leondgarse/keras_cv_attention_models/releases/download/cmt/cmt_tiny_160_imagenet.h5)                     | 9.5M   | 0.65G | 160   | 78.94    | 315.566 qps  |\n  | - [224, (fine-tuned 69 epochs)](https://github.com/leondgarse/keras_cv_attention_models/releases/download/cmt/cmt_tiny_224_imagenet.h5)      | 9.5M   | 1.32G | 224   | 80.73    | 254.87 qps   |\n  | [CMTTiny_torch, (1000 epochs)](https://github.com/leondgarse/keras_cv_attention_models/releases/download/cmt/cmt_tiny_torch_160_imagenet.h5)       | 9.5M   | 0.65G | 160   | 79.2     | 338.207 qps  |\n  | [CMTXS_torch](https://github.com/leondgarse/keras_cv_attention_models/releases/download/cmt/cmt_xs_torch_192_imagenet.h5)                        | 15.2M  | 1.58G | 192   | 81.8     | 241.288 qps  |\n  | [CMTSmall_torch](https://github.com/leondgarse/keras_cv_attention_models/releases/download/cmt/cmt_small_torch_224_imagenet.h5)                     | 25.1M  | 4.09G | 224   | 83.5     | 171.109 qps  |\n  | [CMTBase_torch](https://github.com/leondgarse/keras_cv_attention_models/releases/download/cmt/cmt_base_torch_256_imagenet.h5)                      | 45.7M  | 9.42G | 256   | 84.5     | 103.34 qps   |\n## CoaT\n  - [Keras CoaT](keras_cv_attention_models/coat) is for [PDF 2104.06399 CoaT: Co-Scale Conv-Attentional Image Transformers](http://arxiv.org/abs/2104.06399).\n\n  | Model         | Params | FLOPs | Input | Top1 Acc | T4 Inference |\n  | ------------- | ------ | ----- | ----- | -------- | ------------ |\n  | [CoaTLiteTiny](https://github.com/leondgarse/keras_cv_attention_models/releases/download/coat/coat_lite_tiny_imagenet.h5)  | 5.7M   | 1.60G | 224   | 77.5     | 450.27 qps   |\n  | [CoaTLiteMini](https://github.com/leondgarse/keras_cv_attention_models/releases/download/coat/coat_lite_mini_imagenet.h5)  | 11M    | 2.00G | 224   | 79.1     | 452.884 qps  |\n  | [CoaTLiteSmall](https://github.com/leondgarse/keras_cv_attention_models/releases/download/coat/coat_lite_small_imagenet.h5) | 20M    | 3.97G | 224   | 81.9     | 248.846 qps  |\n  | [CoaTTiny](https://github.com/leondgarse/keras_cv_attention_models/releases/download/coat/coat_tiny_imagenet.h5)      | 5.5M   | 4.33G | 224   | 78.3     | 152.495 qps  |\n  | [CoaTMini](https://github.com/leondgarse/keras_cv_attention_models/releases/download/coat/coat_mini_imagenet.h5)      | 10M    | 6.78G | 224   | 81.0     | 124.845 qps  |\n## CoAtNet\n  - [Keras CoAtNet](keras_cv_attention_models/coatnet) is for [PDF 2106.04803 CoAtNet: Marrying Convolution and Attention for All Data Sizes](https://arxiv.org/pdf/2106.04803.pdf).\n\n  | Model                               | Params | FLOPs  | Input | Top1 Acc | T4 Inference |\n  | ----------------------------------- | ------ | ------ | ----- | -------- | ------------ |\n  | [CoAtNet0, 160, (105 epochs)](https://github.com/leondgarse/keras_cv_attention_models/releases/download/coatnet/coatnet0_160_imagenet.h5) | 23.3M  | 2.09G  | 160   | 80.48    | 584.059 qps  |\n  | [CoAtNet0, (305 epochs)](https://github.com/leondgarse/keras_cv_attention_models/releases/download/coatnet/coatnet0_224_imagenet.h5) | 23.8M  | 4.22G  | 224   | 82.79    | 400.333 qps  |\n  | CoAtNet0                            | 25M    | 4.6G   | 224   | 82.0     | 400.333 qps  |\n  | - use_dw_strides=False              | 25M    | 4.2G   | 224   | 81.6     | 461.197 qps  |\n  | CoAtNet1                            | 42M    | 8.8G   | 224   | 83.5     | 206.954 qps  |\n  | - use_dw_strides=False              | 42M    | 8.4G   | 224   | 83.3     | 228.938 qps  |\n  | CoAtNet2                            | 75M    | 16.6G  | 224   | 84.1     | 156.359 qps  |\n  | - use_dw_strides=False              | 75M    | 15.7G  | 224   | 84.1     | 165.846 qps  |\n  | CoAtNet2, 21k_ft1k                  | 75M    | 16.6G  | 224   | 87.1     | 156.359 qps  |\n  | CoAtNet3                            | 168M   | 34.7G  | 224   | 84.5     | 95.0703 qps  |\n  | CoAtNet3, 21k_ft1k                  | 168M   | 34.7G  | 224   | 87.6     | 95.0703 qps  |\n  | CoAtNet3, 21k_ft1k                  | 168M   | 203.1G | 512   | 87.9     | 95.0703 qps  |\n  | CoAtNet4, 21k_ft1k                  | 275M   | 360.9G | 512   | 88.1     | 74.6022 qps  |\n  | CoAtNet4, 21k_ft1k, PT-RA-E150      | 275M   | 360.9G | 512   | 88.56    | 74.6022 qps  |\n## ConvNeXt\n  - [Keras ConvNeXt](keras_cv_attention_models/convnext) is for [PDF 2201.03545 A ConvNet for the 2020s](https://arxiv.org/pdf/2201.03545.pdf).\n\n  | Model                   | Params | FLOPs   | Input | Top1 Acc | T4 Inference |\n  | ----------------------- | ------ | ------- | ----- | -------- | ------------ |\n  | [ConvNeXtTiny](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_tiny_imagenet.h5)            | 28M    | 4.49G   | 224   | 82.1     | 361.58 qps   |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_tiny_224_imagenet21k-ft1k.h5)      | 28M    | 4.49G   | 224   | 82.9     | 361.58 qps   |\n  | - [21k_ft1k, 384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_tiny_384_imagenet21k-ft1k.h5) | 28M    | 13.19G  | 384   | 84.1     | 182.134 qps  |\n  | [ConvNeXtSmall](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_small_imagenet.h5)           | 50M    | 8.73G   | 224   | 83.1     | 202.007 qps  |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_small_224_imagenet21k-ft1k.h5)      | 50M    | 8.73G   | 224   | 84.6     | 202.007 qps  |\n  | - [21k_ft1k, 384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_small_384_imagenet21k-ft1k.h5) | 50M    | 25.67G  | 384   | 85.8     | 108.125 qps  |\n  | [ConvNeXtBase](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_base_224_imagenet.h5)            | 89M    | 15.42G  | 224   | 83.8     | 160.036 qps  |\n  | - [384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_base_384_imagenet.h5)                   | 89M    | 45.32G  | 384   | 85.1     | 83.3095 qps  |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_base_224_imagenet21k-ft1k.h5)      | 89M    | 15.42G  | 224   | 85.8     | 160.036 qps  |\n  | - [21k_ft1k, 384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_base_384_imagenet21k-ft1k.h5) | 89M    | 45.32G  | 384   | 86.8     | 83.3095 qps  |\n  | [ConvNeXtLarge](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_large_224_imagenet.h5)           | 198M   | 34.46G  | 224   | 84.3     | 102.27 qps   |\n  | - [384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_large_384_imagenet.h5)                   | 198M   | 101.28G | 384   | 85.5     | 47.2086 qps  |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_large_224_imagenet21k-ft1k.h5)      | 198M   | 34.46G  | 224   | 86.6     | 102.27 qps   |\n  | - [21k_ft1k, 384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_large_384_imagenet21k-ft1k.h5) | 198M   | 101.28G | 384   | 87.5     | 47.2086 qps  |\n  | [ConvNeXtXlarge, 21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_xlarge_224_imagenet21k-ft1k.h5)     | 350M   | 61.06G  | 224   | 87.0     | 40.5776 qps  |\n  | - [21k_ft1k, 384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_xlarge_384_imagenet21k-ft1k.h5)              | 350M   | 179.43G | 384   | 87.8     | 21.797 qps   |\n  | [ConvNeXtXXLarge, clip](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_xxlarge_clip-ft1k.h5)   | 846M   | 198.09G | 256   | 88.6     |              |\n## ConvNeXtV2\n  - [Keras ConvNeXt](keras_cv_attention_models/convnext) includes implementation of [PDF 2301.00808 ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders](https://arxiv.org/pdf/2301.00808.pdf). **Please note the CC-BY-NC 4.0 license on theses weights, non-commercial use only**.\n\n  | Model                   | Params | FLOPs  | Input | Top1 Acc | T4 Inference |\n  | ----------------------- | ------ | ------ | ----- | -------- | ------------ |\n  | [ConvNeXtV2Atto](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_v2_atto_imagenet.h5)          | 3.7M   | 0.55G  | 224   | 76.7     | 705.822 qps  |\n  | [ConvNeXtV2Femto](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_v2_femto_imagenet.h5)         | 5.2M   | 0.78G  | 224   | 78.5     | 728.02 qps   |\n  | [ConvNeXtV2Pico](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_v2_pico_imagenet.h5)          | 9.1M   | 1.37G  | 224   | 80.3     | 591.502 qps  |\n  | [ConvNeXtV2Nano](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_v2_nano_imagenet.h5)          | 15.6M  | 2.45G  | 224   | 81.9     | 471.918 qps  |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_v2_nano_224_imagenet21k-ft1k.h5)      | 15.6M  | 2.45G  | 224   | 82.1     | 471.918 qps  |\n  | - [21k_ft1k, 384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_v2_nano_384_imagenet21k-ft1k.h5) | 15.6M  | 7.21G  | 384   | 83.4     | 213.802 qps  |\n  | [ConvNeXtV2Tiny](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_v2_tiny_imagenet.h5)          | 28.6M  | 4.47G  | 224   | 83.0     | 301.982 qps  |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_v2_tiny_224_imagenet21k-ft1k.h5)      | 28.6M  | 4.47G  | 224   | 83.9     | 301.982 qps  |\n  | - [21k_ft1k, 384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_v2_tiny_384_imagenet21k-ft1k.h5) | 28.6M  | 13.1G  | 384   | 85.1     | 139.578 qps  |\n  | [ConvNeXtV2Base](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_v2_base_imagenet.h5)          | 89M    | 15.4G  | 224   | 84.9     | 132.575 qps  |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_v2_base_224_imagenet21k-ft1k.h5)      | 89M    | 15.4G  | 224   | 86.8     | 132.575 qps  |\n  | - [21k_ft1k, 384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_v2_base_384_imagenet21k-ft1k.h5) | 89M    | 45.2G  | 384   | 87.7     | 66.5729 qps  |\n  | [ConvNeXtV2Large](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_v2_large_imagenet.h5)         | 198M   | 34.4G  | 224   | 85.8     | 86.8846 qps  |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_v2_large_224_imagenet21k-ft1k.h5)      | 198M   | 34.4G  | 224   | 87.3     | 86.8846 qps  |\n  | - [21k_ft1k, 384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_v2_large_384_imagenet21k-ft1k.h5) | 198M   | 101.1G | 384   | 88.2     | 24.4542 qps  |\n  | [ConvNeXtV2Huge](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_v2_huge_imagenet.h5)          | 660M   | 115G   | 224   | 86.3     |              |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_v2_huge_384_imagenet21k-ft1k.h5)      | 660M   | 337.9G | 384   | 88.7     |              |\n  | - [21k_ft1k, 384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/convnext/convnext_v2_huge_512_imagenet21k-ft1k.h5) | 660M   | 600.8G | 512   | 88.9     |              |\n## CoTNet\n  - [Keras CoTNet](keras_cv_attention_models/cotnet) is for [PDF 2107.12292 Contextual Transformer Networks for Visual Recognition](https://arxiv.org/pdf/2107.12292.pdf).\n\n  | Model        | Params | FLOPs  | Input | Top1 Acc | T4 Inference |\n  | ------------ |:------:| ------ | ----- |:--------:| ------------ |\n  | [CotNet50](https://github.com/leondgarse/keras_cv_attention_models/releases/download/cotnet/cotnet50_224_imagenet.h5)     | 22.2M  | 3.25G  | 224   |   81.3   | 324.913 qps  |\n  | [CotNetSE50D](https://github.com/leondgarse/keras_cv_attention_models/releases/download/cotnet/cotnet_se50d_224_imagenet.h5)  | 23.1M  | 4.05G  | 224   |   81.6   | 513.077 qps  |\n  | [CotNet101](https://github.com/leondgarse/keras_cv_attention_models/releases/download/cotnet/cotnet101_224_imagenet.h5)    | 38.3M  | 6.07G  | 224   |   82.8   | 183.824 qps  |\n  | [CotNetSE101D](https://github.com/leondgarse/keras_cv_attention_models/releases/download/cotnet/cotnet_se101d_224_imagenet.h5) | 40.9M  | 8.44G  | 224   |   83.2   | 251.487 qps  |\n  | [CotNetSE152D](https://github.com/leondgarse/keras_cv_attention_models/releases/download/cotnet/cotnet_se152d_224_imagenet.h5) | 55.8M  | 12.22G | 224   |   84.0   | 175.469 qps  |\n  | [CotNetSE152D](https://github.com/leondgarse/keras_cv_attention_models/releases/download/cotnet/cotnet_se152d_320_imagenet.h5) | 55.8M  | 24.92G | 320   |   84.6   | 175.469 qps  |\n## CSPNeXt\n  - [Keras CSPNeXt](keras_cv_attention_models/cspnext) is for backbone of [PDF 2212.07784 RTMDet: An Empirical Study of Designing Real-Time Object Detectors](https://arxiv.org/abs/2212.07784).\n\n  | Model         | Params | FLOPs | Input | Top1 Acc | T4 Inference |\n  | ------------- | ------ | ----- | ----- | -------- | -------- |\n  | [CSPNeXtTiny](https://github.com/leondgarse/keras_cv_attention_models/releases/download/cspnext/cspnext_tiny_imagenet.h5)   | 2.73M  | 0.34G | 224   | 69.44    |  |\n  | [CSPNeXtSmall](https://github.com/leondgarse/keras_cv_attention_models/releases/download/cspnext/cspnext_small_imagenet.h5)  | 4.89M  | 0.66G | 224   | 74.41    |  |\n  | [CSPNeXtMedium](https://github.com/leondgarse/keras_cv_attention_models/releases/download/cspnext/cspnext_medium_imagenet.h5) | 13.05M | 1.92G | 224   | 79.27    |  |\n  | [CSPNeXtLarge](https://github.com/leondgarse/keras_cv_attention_models/releases/download/cspnext/cspnext_large_imagenet.h5)  | 27.16M | 4.19G | 224   | 81.30    |  |\n  | [CSPNeXtXLarge](https://github.com/leondgarse/keras_cv_attention_models/releases/download/cspnext/cspnext_xlarge_imagenet.h5) | 48.85M | 7.75G | 224   | 82.10    |  |\n## DaViT\n  - [Keras DaViT](keras_cv_attention_models/davit) is for [PDF 2204.03645 DaViT: Dual Attention Vision Transformers](https://arxiv.org/pdf/2204.03645.pdf).\n\n  | Model              | Params | FLOPs  | Input | Top1 Acc | T4 Inference |\n  | ------------------ | ------ | ------ | ----- | -------- | ------------ |\n  | [DaViT_T](https://github.com/leondgarse/keras_cv_attention_models/releases/download/davit/davit_t_imagenet.h5)            | 28.36M | 4.56G  | 224   | 82.8     | 224.563 qps  |\n  | [DaViT_S](https://github.com/leondgarse/keras_cv_attention_models/releases/download/davit/davit_s_imagenet.h5)            | 49.75M | 8.83G  | 224   | 84.2     | 145.838 qps  |\n  | [DaViT_B](https://github.com/leondgarse/keras_cv_attention_models/releases/download/davit/davit_b_imagenet.h5)            | 87.95M | 15.55G | 224   | 84.6     | 114.527 qps  |\n  | DaViT_L, 21k_ft1k  | 196.8M | 103.2G | 384   | 87.5     | 34.7015 qps  |\n  | DaViT_H, 1.5B      | 348.9M | 327.3G | 512   | 90.2     | 12.363 qps   |\n  | DaViT_G, 1.5B      | 1.406B | 1.022T | 512   | 90.4     |              |\n## DiNAT\n  - [Keras DiNAT](keras_cv_attention_models/nat) is for [PDF 2209.15001 Dilated Neighborhood Attention Transformer](https://arxiv.org/pdf/2209.15001.pdf).\n\n  | Model                     | Params | FLOPs  | Input | Top1 Acc | T4 Inference |\n  | ------------------------- | ------ | ------ | ----- | -------- | ------------ |\n  | [DiNAT_Mini](https://github.com/leondgarse/keras_cv_attention_models/releases/download/nat/dinat_mini_imagenet.h5)                | 20.0M  | 2.73G  | 224   | 81.8     | 83.9943 qps  |\n  | [DiNAT_Tiny](https://github.com/leondgarse/keras_cv_attention_models/releases/download/nat/dinat_tiny_imagenet.h5)                | 27.9M  | 4.34G  | 224   | 82.7     | 61.1902 qps  |\n  | [DiNAT_Small](https://github.com/leondgarse/keras_cv_attention_models/releases/download/nat/dinat_small_imagenet.h5)               | 50.7M  | 7.84G  | 224   | 83.8     | 41.0343 qps  |\n  | [DiNAT_Base](https://github.com/leondgarse/keras_cv_attention_models/releases/download/nat/dinat_base_imagenet.h5)                | 89.8M  | 13.76G | 224   | 84.4     | 30.1332 qps  |\n  | [DiNAT_Large, 21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/nat/dinat_large_224_imagenet21k-ft1k.h5)     | 200.9M | 30.58G | 224   | 86.6     | 18.4936 qps  |\n  | - [21k, (num_classes=21841)](https://github.com/leondgarse/keras_cv_attention_models/releases/download/nat/dinat_large_imagenet21k.h5)   | 200.9M | 30.58G | 224   |          |              |\n  | - [21k_ft1k, 384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/nat/dinat_large_384_imagenet21k-ft1k.h5)           | 200.9M | 89.86G | 384   | 87.4     |              |\n  | [DiNAT_Large_K11, 21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/nat/dinat_large_k11_imagenet21k-ft1k.h5) | 201.1M | 92.57G | 384   | 87.5     |              |\n## DINOv2\n  - [Keras DINOv2](keras_cv_attention_models/beit) includes models from [PDF 2304.07193 DINOv2: Learning Robust Visual Features without Supervision](https://arxiv.org/pdf/2304.07193.pdf).\n\n  | Model              | Params  | FLOPs   | Input | Top1 Acc | T4 Inference |\n  | ------------------ | ------- | ------- | ----- | -------- | ------------ |\n  | [DINOv2_ViT_Small14](https://github.com/leondgarse/keras_cv_attention_models/releases/download/beit/dinov2_vit_small14_518_imagenet.h5) | 22.83M  | 47.23G  | 518   | 81.1     | 165.271 qps  |\n  | [DINOv2_ViT_Base14](https://github.com/leondgarse/keras_cv_attention_models/releases/download/beit/dinov2_vit_base14_518_imagenet.h5)  | 88.12M  | 152.6G  | 518   | 84.5     | 54.9769 qps  |\n  | [DINOv2_ViT_Large14](https://github.com/leondgarse/keras_cv_attention_models/releases/download/beit/dinov2_vit_large14_518_imagenet.h5) | 306.4M  | 509.6G  | 518   | 86.3     | 17.4108 qps  |\n  | [DINOv2_ViT_Giant14](https://github.com/leondgarse/keras_cv_attention_models/releases/download/beit/dinov2_vit_giant14_518_imagenet.h5) | 1139.6M | 1790.3G | 518   | 86.5     |              |\n## EdgeNeXt\n  - [Keras EdgeNeXt](keras_cv_attention_models/edgenext) is for [PDF 2206.10589 EdgeNeXt: Efficiently Amalgamated CNN-Transformer Architecture for Mobile Vision Applications](https://arxiv.org/pdf/2206.10589.pdf).\n\n  | Model             | Params | FLOPs  | Input | Top1 Acc | T4 Inference |\n  | ----------------- | ------ | ------ | ----- | -------- | ------------ |\n  | [EdgeNeXt_XX_Small](https://github.com/leondgarse/keras_cv_attention_models/releases/download/edgenext/edgenext_xx_small_256_imagenet.h5) | 1.33M  | 266M   | 256   | 71.23    | 902.957 qps  |\n  | [EdgeNeXt_X_Small](https://github.com/leondgarse/keras_cv_attention_models/releases/download/edgenext/edgenext_x_small_256_imagenet.h5)  | 2.34M  | 547M   | 256   | 74.96    | 638.346 qps  |\n  | [EdgeNeXt_Small](https://github.com/leondgarse/keras_cv_attention_models/releases/download/edgenext/edgenext_small_256_imagenet.h5)    | 5.59M  | 1.27G  | 256   | 79.41    | 536.762 qps  |\n  | - [usi](https://github.com/leondgarse/keras_cv_attention_models/releases/download/edgenext/edgenext_small_256_usi.h5)             | 5.59M  | 1.27G  | 256   | 81.07    | 536.762 qps  |\n  | [EdgeNeXt_Base](https://github.com/leondgarse/keras_cv_attention_models/releases/download/edgenext/edgenext_base_256_imagenet.h5)     | 18.5M  | 3.86G  | 256   | 82.47    | 383.461 qps  |\n  | - [usi](https://github.com/leondgarse/keras_cv_attention_models/releases/download/edgenext/edgenext_base_256_usi.h5)             | 18.5M  | 3.86G  | 256   | 83.31    | 383.461 qps  |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/edgenext/edgenext_base_256_imagenet-ft1k.h5)        | 18.5M  | 3.86G  | 256   | 83.68    | 383.461 qps  |\n## EfficientFormer\n  - [Keras EfficientFormer](keras_cv_attention_models/efficientformer) is for [PDF 2206.01191 EfficientFormer: Vision Transformers at MobileNet Speed](https://arxiv.org/pdf/2206.01191.pdf).\n\n  | Model                      | Params | FLOPs | Input | Top1 Acc | T4 Inference |\n  | -------------------------- | ------ | ----- | ----- | -------- | ------------ |\n  | [EfficientFormerL1, distill](https://github.com/leondgarse/keras_cv_attention_models/releases/download/levit/efficientformer_l1_224_imagenet.h5) | 12.3M  | 1.31G | 224   | 79.2     | 1214.22 qps  |\n  | [EfficientFormerL3, distill](https://github.com/leondgarse/keras_cv_attention_models/releases/download/levit/efficientformer_l3_224_imagenet.h5) | 31.4M  | 3.95G | 224   | 82.4     | 596.705 qps  |\n  | [EfficientFormerL7, distill](https://github.com/leondgarse/keras_cv_attention_models/releases/download/levit/efficientformer_l7_224_imagenet.h5) | 74.4M  | 9.79G | 224   | 83.3     | 298.434 qps  |\n## EfficientFormerV2\n  - [Keras EfficientFormer](keras_cv_attention_models/efficientformer) includes implementation of [PDF 2212.08059 Rethinking Vision Transformers for MobileNet Size and Speed](https://arxiv.org/pdf/2212.08059.pdf).\n\n  | Model                        | Params | FLOPs  | Input | Top1 Acc | T4 Inference |\n  | ---------------------------- | ------ | ------ | ----- | -------- | ------------ |\n  | [EfficientFormerV2S0, distill](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientformer/efficientformer_v2_s0_224_imagenet.h5) | 3.60M  | 405.2M | 224   | 76.2     | 1114.38 qps  |\n  | [EfficientFormerV2S1, distill](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientformer/efficientformer_v2_s1_224_imagenet.h5) | 6.19M  | 665.6M | 224   | 79.7     | 841.186 qps  |\n  | [EfficientFormerV2S2, distill](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientformer/efficientformer_v2_s2_224_imagenet.h5) | 12.7M  | 1.27G  | 224   | 82.0     | 573.9 qps    |\n  | [EfficientFormerV2L, distill](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientformer/efficientformer_v2_l_224_imagenet.h5)  | 26.3M  | 2.59G  | 224   | 83.5     | 377.224 qps  |\n## EfficientNet\n  - [Keras EfficientNet](keras_cv_attention_models/efficientnet) includes implementation of [PDF 1911.04252 Self-training with Noisy Student improves ImageNet classification](https://arxiv.org/pdf/1911.04252.pdf).\n\n  | Model                          | Params | FLOPs   | Input | Top1 Acc | T4 Inference |\n  | ------------------------------ | ------ | ------- | ----- | -------- | ------------ |\n  | [EfficientNetV1B0](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv1_pretrained/efficientnetv1-b0-imagenet.h5)               | 5.3M   | 0.39G   | 224   | 77.6     | 1129.93 qps  |\n  | - [NoisyStudent](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv1_pretrained/efficientnetv1-b0-noisy_student.h5)                 | 5.3M   | 0.39G   | 224   | 78.8     | 1129.93 qps  |\n  | [EfficientNetV1B1](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv1_pretrained/efficientnetv1-b1-imagenet.h5)               | 7.8M   | 0.70G   | 240   | 79.6     | 758.639 qps  |\n  | - [NoisyStudent](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv1_pretrained/efficientnetv1-b1-noisy_student.h5)                 | 7.8M   | 0.70G   | 240   | 81.5     | 758.639 qps  |\n  | [EfficientNetV1B2](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv1_pretrained/efficientnetv1-b2-imagenet.h5)               | 9.1M   | 1.01G   | 260   | 80.5     | 668.959 qps  |\n  | - [NoisyStudent](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv1_pretrained/efficientnetv1-b2-noisy_student.h5)                 | 9.1M   | 1.01G   | 260   | 82.4     | 668.959 qps  |\n  | [EfficientNetV1B3](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv1_pretrained/efficientnetv1-b3-imagenet.h5)               | 12.2M  | 1.86G   | 300   | 81.9     | 473.607 qps  |\n  | - [NoisyStudent](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv1_pretrained/efficientnetv1-b3-noisy_student.h5)                 | 12.2M  | 1.86G   | 300   | 84.1     | 473.607 qps  |\n  | [EfficientNetV1B4](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv1_pretrained/efficientnetv1-b4-imagenet.h5)               | 19.3M  | 4.46G   | 380   | 83.3     | 265.244 qps  |\n  | - [NoisyStudent](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv1_pretrained/efficientnetv1-b4-noisy_student.h5)                 | 19.3M  | 4.46G   | 380   | 85.3     | 265.244 qps  |\n  | [EfficientNetV1B5](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv1_pretrained/efficientnetv1-b5-imagenet.h5)               | 30.4M  | 10.40G  | 456   | 84.3     | 146.758 qps  |\n  | - [NoisyStudent](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv1_pretrained/efficientnetv1-b5-noisy_student.h5)                 | 30.4M  | 10.40G  | 456   | 86.1     | 146.758 qps  |\n  | [EfficientNetV1B6](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv1_pretrained/efficientnetv1-b6-imagenet.h5)               | 43.0M  | 19.29G  | 528   | 84.8     | 88.0369 qps  |\n  | - [NoisyStudent](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv1_pretrained/efficientnetv1-b6-noisy_student.h5)                 | 43.0M  | 19.29G  | 528   | 86.4     | 88.0369 qps  |\n  | [EfficientNetV1B7](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv1_pretrained/efficientnetv1-b7-imagenet.h5)               | 66.3M  | 38.13G  | 600   | 85.2     | 52.6616 qps  |\n  | - [NoisyStudent](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv1_pretrained/efficientnetv1-b7-noisy_student.h5)                 | 66.3M  | 38.13G  | 600   | 86.9     | 52.6616 qps  |\n  | [EfficientNetV1L2, NoisyStudent](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv1_pretrained/efficientnetv1-l2-noisy_student.h5) | 480.3M | 477.98G | 800   | 88.4     |              |\n## EfficientNetEdgeTPU\n  - [Keras EfficientNetEdgeTPU](keras_cv_attention_models/efficientnet) includes implementation of [PDF 1911.04252 Self-training with Noisy Student improves ImageNet classification](https://arxiv.org/pdf/1911.04252.pdf).\n\n  | Model                          | Params | FLOPs   | Input | Top1 Acc | T4 Inference |\n  | ------------------------------ | ------ | ------- | ----- | -------- | ------------ |\n  | [EfficientNetEdgeTPUSmall](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv1_pretrained/efficientnetedgetpu-small-imagenet.h5)       | 5.49M  | 1.79G   | 224   | 78.07    | 1459.38 qps  |\n  | [EfficientNetEdgeTPUMedium](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv1_pretrained/efficientnetedgetpu-medium-imagenet.h5)      | 6.90M  | 3.01G   | 240   | 79.25    | 1028.95 qps  |\n  | [EfficientNetEdgeTPULarge](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv1_pretrained/efficientnetedgetpu-large-imagenet.h5)       | 10.59M | 7.94G   | 300   | 81.32    | 527.034 qps  |\n## EfficientNetV2\n  - [Keras EfficientNet](keras_cv_attention_models/efficientnet) includes implementation of [PDF 2104.00298 EfficientNetV2: Smaller Models and Faster Training](https://arxiv.org/abs/2104.00298).\n\n  | Model                      | Params | FLOPs  | Input | Top1 Acc | T4 Inference |\n  | -------------------------- | ------ | ------ | ----- | -------- | ------------ |\n  | [EfficientNetV2B0](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv2_pretrained/efficientnetv2-b0-imagenet.h5)           | 7.1M   | 0.72G  | 224   | 78.7     | 1109.84 qps  |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv2_pretrained/efficientnetv2-b0-21k-ft1k.h5)         | 7.1M   | 0.72G  | 224   | 77.55?   | 1109.84 qps  |\n  | [EfficientNetV2B1](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv2_pretrained/efficientnetv2-b1-imagenet.h5)           | 8.1M   | 1.21G  | 240   | 79.8     | 842.372 qps  |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv2_pretrained/efficientnetv2-b1-21k-ft1k.h5)         | 8.1M   | 1.21G  | 240   | 79.03?   | 842.372 qps  |\n  | [EfficientNetV2B2](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv2_pretrained/efficientnetv2-b2-imagenet.h5)           | 10.1M  | 1.71G  | 260   | 80.5     | 762.865 qps  |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv2_pretrained/efficientnetv2-b2-21k-ft1k.h5)         | 10.1M  | 1.71G  | 260   | 79.48?   | 762.865 qps  |\n  | [EfficientNetV2B3](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv2_pretrained/efficientnetv2-b3-imagenet.h5)           | 14.4M  | 3.03G  | 300   | 82.1     | 548.501 qps  |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv2_pretrained/efficientnetv2-b3-21k-ft1k.h5)         | 14.4M  | 3.03G  | 300   | 82.46?   | 548.501 qps  |\n  | [EfficientNetV2T](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv2_pretrained/efficientnetv2-t-imagenet.h5)            | 13.6M  | 3.18G  | 288   | 82.34    | 496.483 qps  |\n  | [EfficientNetV2T_GC](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv2_pretrained/efficientnetv2-t-gc-imagenet.h5)         | 13.7M  | 3.19G  | 288   | 82.46    | 368.763 qps  |\n  | [EfficientNetV2S](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv2_pretrained/efficientnetv2-s-imagenet.h5)            | 21.5M  | 8.41G  | 384   | 83.9     | 344.109 qps  |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv2_pretrained/efficientnetv2-s-21k-ft1k.h5)         | 21.5M  | 8.41G  | 384   | 84.9     | 344.109 qps  |\n  | [EfficientNetV2M](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv2_pretrained/efficientnetv2-m-imagenet.h5)            | 54.1M  | 24.69G | 480   | 85.2     | 145.346 qps  |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv2_pretrained/efficientnetv2-m-21k-ft1k.h5)         | 54.1M  | 24.69G | 480   | 86.2     | 145.346 qps  |\n  | [EfficientNetV2L](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv2_pretrained/efficientnetv2-l-imagenet.h5)            | 119.5M | 56.27G | 480   | 85.7     | 85.6514 qps  |\n  | - [21k_ft1k](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv2_pretrained/efficientnetv2-l-21k-ft1k.h5)         | 119.5M | 56.27G | 480   | 86.9     | 85.6514 qps  |\n  | [EfficientNetV2XL, 21k_ft1k](https://github.com/leondgarse/keras_efficientnet_v2/releases/download/effnetv2_pretrained/efficientnetv2-xl-21k-ft1k.h5) | 206.8M | 93.66G | 512   | 87.2     | 55.141 qps   |\n## EfficientViT_B\n  - [Keras EfficientViT_B](keras_cv_attention_models/efficientvit) is for Paper [PDF 2205.14756 EfficientViT: Lightweight Multi-Scale Attention for On-Device Semantic Segmentation](https://arxiv.org/pdf/2205.14756.pdf).\n\n  | Model           | Params | FLOPs | Input | Top1 Acc | T4 Inference |\n  | --------------- | ------ | ----- | ----- | -------- | ------------ |\n  | [EfficientViT_B0](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientvit/efficientvit_b0_224_imagenet.h5) | 3.41M  | 0.12G | 224   | 71.6 ?   | 1581.76 qps  |\n  | [EfficientViT_B1](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientvit/efficientvit_b1_224_imagenet.h5) | 9.10M  | 0.58G | 224   | 79.4     | 943.587 qps  |\n  | - [256](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientvit/efficientvit_b1_256_imagenet.h5)           | 9.10M  | 0.78G | 256   | 79.9     | 840.844 qps  |\n  | - [288](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientvit/efficientvit_b1_288_imagenet.h5)            | 9.10M  | 1.03G | 288   | 80.4     | 680.088 qps  |\n  | [EfficientViT_B2](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientvit/efficientvit_b2_224_imagenet.h5) | 24.33M | 1.68G | 224   | 82.1     | 583.295 qps  |\n  | - [256](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientvit/efficientvit_b2_256_imagenet.h5)            | 24.33M | 2.25G | 256   | 82.7     | 507.187 qps  |\n  | - [288](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientvit/efficientvit_b2_288_imagenet.h5)            | 24.33M | 2.92G | 288   | 83.1     | 419.93 qps   |\n  | [EfficientViT_B3](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientvit/efficientvit_b3_224_imagenet.h5) | 48.65M | 4.14G | 224   | 83.5     | 329.764 qps  |\n  | - [256](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientvit/efficientvit_b3_256_imagenet.h5)            | 48.65M | 5.51G | 256   | 83.8     | 288.605 qps  |\n  | - [288](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientvit/efficientvit_b3_288_imagenet.h5)            | 48.65M | 7.14G | 288   | 84.2     | 229.992 qps  |\n  | [EfficientViT_L1](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientvit/efficientvit_l1_224_imagenet.h5) | 52.65M | 5.28G | 224   | 84.48    | 503.068 qps |\n  | [EfficientViT_L2](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientvit/efficientvit_l2_224_imagenet.h5) | 63.71M | 6.98G | 224   | 85.05    | 396.255 qps |\n  | - [384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientvit/efficientvit_l2_384_imagenet.h5)            | 63.71M | 20.7G | 384   | 85.98    | 207.322 qps |\n  | [EfficientViT_L3](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientvit/efficientvit_l3_224_imagenet.h5) | 246.0M | 27.6G | 224   | 85.814   | 174.926 qps |\n  | - [384](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientvit/efficientvit_l3_384_imagenet.h5)            | 246.0M | 81.6G | 384   | 86.408   | 86.895 qps  |\n## EfficientViT_M\n  - [Keras EfficientViT_M](keras_cv_attention_models/efficientvit) is for Paper [PDF 2305.07027 EfficientViT: Memory Efficient Vision Transformer with Cascaded Group Attention](https://arxiv.org/pdf/2305.07027.pdf).\n\n  | Model           | Params | FLOPs | Input | Top1 Acc | T4 Inference |\n  | --------------- | ------ | ----- | ----- | -------- | ------------ |\n  | [EfficientViT_M0](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientvit/efficientvit_m0_224_imagenet.h5) | 2.35M  | 79.4M | 224   | 63.2     | 814.522 qps  |\n  | [EfficientViT_M1](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientvit/efficientvit_m1_224_imagenet.h5) | 2.98M  | 167M  | 224   | 68.4     | 948.041 qps  |\n  | [EfficientViT_M2](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientvit/efficientvit_m2_224_imagenet.h5) | 4.19M  | 201M  | 224   | 70.8     | 906.286 qps  |\n  | [EfficientViT_M3](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientvit/efficientvit_m3_224_imagenet.h5) | 6.90M  | 263M  | 224   | 73.4     | 758.086 qps  |\n  | [EfficientViT_M4](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientvit/efficientvit_m4_224_imagenet.h5) | 8.80M  | 299M  | 224   | 74.3     | 672.891 qps  |\n  | [EfficientViT_M5](https://github.com/leondgarse/keras_cv_attention_models/releases/download/efficientvit/efficientvit_m5_224_imagenet.h5) | 12.47M | 522M  | 224   | 77.1     | 577.254 qps  |\n## EVA\n  - [Keras EVA](keras_cv_attention_models/beit) includes models from [PDF 2211.07636 EVA: Exploring the Limits of Masked Visual Representation Learning at Scale](https://arxiv.org/pdf/2211.07636.pdf).\n\n  | Model                 | Params  | FLOPs    | Input | Top1 Acc | T4 Inference |\n  | --------------------- | ------- | -------- | ----- | -------- | ------------ |\n  | [EvaLargePatch14, 21k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/beit/eva_large_patch14_196_imagenet21k-ft1k.h5)  | 304.14M | 61.65G   | 196   | 88.59    | 115.532 qps  |\n  | - [21k_ft1k, 336](https://github.com/leondgarse/keras_cv_attention_models/releases/download/beit/eva_large_patch14_336_imagenet21k-ft1k.h5)            | 304.53M | 191.55G  | 336   | 89.20    | 53.3467 qps  |\n  | [EvaGiantPatch14, clip](https://github.com/leondgarse/keras_cv_attention_models/releases/download/beit/eva_giant_patch14_224_imagenet21k-ft1k.h5) | 1012.6M | 267.40G  | 224   | 89.10    |              |\n  | - [m30m](https://github.com/leondgarse/keras_cv_attention_models/releases/download/beit/eva_giant_patch14_336_imagenet21k-ft1k.h5)                | 1013.0M | 621.45G  | 336   | 89.57    |              |\n  | - [m30m](https://github.com/leondgarse/keras_cv_attention_models/releases/download/beit/eva_giant_patch14_560_imagenet21k-ft1k.h5)                | 1014.4M | 1911.61G | 560   | 89.80    |              |\n## EVA02\n  - [Keras EVA02](keras_cv_attention_models/beit) includes models from [PDF 2303.11331 EVA: EVA-02: A Visual Representation for Neon Genesis](https://arxiv.org/pdf/2303.11331.pdf).\n\n  | Model                                  | Params  | FLOPs   | Input | Top1 Acc | T4 Inference |\n  | -------------------------------------- | ------- | ------- | ----- | -------- | ------------ |\n  | [EVA02TinyPatch14, mim_in22k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/beit/eva02_tiny_patch14_336_mim_in22k_ft1k.h5)       | 5.76M   | 4.72G   | 336   | 80.658   | 320.123 qps  |\n  | [EVA02SmallPatch14, mim_in22k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/beit/eva02_small_patch14_336_mim_in22k_ft1k.h5)      | 22.13M  | 15.57G  | 336   | 85.74    | 161.774 qps  |\n  | [EVA02BasePatch14, mim_in22k_ft22k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/beit/eva02_base_patch14_448_mim_in22k_ft22k_ft1k.h5) | 87.12M  | 107.6G  | 448   | 88.692   | 34.3962 qps  |\n  | [EVA02LargePatch14, mim_m38m_ft22k_ft1k](https://github.com/leondgarse/keras_cv_attention_models/releases/download/beit/eva02_large_patch14_448_mim_m38m_ft22k_ft1k.h5) | 305.08M | 363.68G | 448   | 90.054   |              |\n## FasterNet\n  - [Keras FasterNet](keras_cv_attention_models/fasternet) includes implementation of [PDF 2303.03667 Run, Don’t Walk: Chasing Higher FLOPS for Faster Neural Networks ](https://arxiv.org/pdf/2303.03667.pdf).\n\n  | Model       | Params | FLOPs  | Input | Top1 Acc | T4 Inference |\n  | ----------- | ------ | ------ | ----- | -------- | ------------ |\n  | [FasterNetT0](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fasternet/fasternet_t0_imagenet.h5) | 3.9M   | 0.34G  | 224   | 71.9     | 1890.83 qps  |\n  | [FasterNetT1](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fasternet/fasternet_t1_imagenet.h5) | 7.6M   | 0.85G  | 224   | 76.2     | 1788.16 qps  |\n  | [FasterNetT2](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fasternet/fasternet_t2_imagenet.h5) | 15.0M  | 1.90G  | 224   | 78.9     | 1353.12 qps  |\n  | [FasterNetS](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fasternet/fasternet_s_imagenet.h5)  | 31.1M  | 4.55G  | 224   | 81.3     | 818.814 qps  |\n  | [FasterNetM](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fasternet/fasternet_m_imagenet.h5)  | 53.5M  | 8.72G  | 224   | 83.0     | 436.383 qps  |\n  | [FasterNetL](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fasternet/fasternet_l_imagenet.h5)  | 93.4M  | 15.49G | 224   | 83.5     | 319.809 qps  |\n## FasterViT\n  - [Keras FasterViT](keras_cv_attention_models/fastervit) includes implementation of [PDF 2306.06189 FasterViT: Fast Vision Transformers with Hierarchical Attention](https://arxiv.org/pdf/2306.06189.pdf).\n\n  | Model      | Params   | FLOPs   | Input | Top1 Acc | T4 Inference |\n  | ---------- | -------- | ------- | ----- | -------- | ------------ |\n  | [FasterViT0](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fastervit/fastervit_0_224_imagenet.h5) | 31.40M   | 3.51G   | 224   | 82.1     | 716.809 qps  |\n  | [FasterViT1](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fastervit/fastervit_1_224_imagenet.h5) | 53.37M   | 5.52G   | 224   | 83.2     | 491.971 qps  |\n  | [FasterViT2](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fastervit/fastervit_2_224_imagenet.h5) | 75.92M   | 9.00G   | 224   | 84.2     | 377.006 qps  |\n  | [FasterViT3](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fastervit/fastervit_3_224_imagenet.h5) | 159.55M  | 18.75G  | 224   | 84.9     | 216.481 qps  |\n  | [FasterViT4](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fastervit/fastervit_4_224_imagenet.h5) | 351.12M  | 41.57G  | 224   | 85.4     | 71.6303 qps  |\n  | [FasterViT5](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fastervit/fastervit_5_224_imagenet.h5) | 957.52M  | 114.08G | 224   | 85.6     |              |\n  | [FasterViT6](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fastervit/fastervit_6_224_imagenet.1.h5), [+.2](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fastervit/fastervit_6_224_imagenet.2.h5) | 1360.33M | 144.13G | 224   | 85.8     |              |\n## FastViT\n  - [Keras FastViT](keras_cv_attention_models/fastvit) includes implementation of [PDF 2303.14189 FastViT: A Fast Hybrid Vision Transformer using Structural Reparameterization](https://arxiv.org/pdf/2303.14189.pdf).\n\n  | Model         | Params | FLOPs | Input | Top1 Acc | T4 Inference |\n  | ------------- | ------ | ----- | ----- | -------- | ------------ |\n  | [FastViT_T8](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fastvit/fastvit_t8_imagenet.h5)     | 4.03M  | 0.65G | 256   | 76.2     | 1020.29 qps  |\n  | - [distill](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fastvit/fastvit_t8_distill.h5)       | 4.03M  | 0.65G | 256   | 77.2     | 1020.29 qps  |\n  | - deploy=True | 3.99M  | 0.64G | 256   | 76.2     | 1323.14 qps  |\n  | [FastViT_T12](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fastvit/fastvit_t12_imagenet.h5)   | 7.55M  | 1.34G | 256   | 79.3     | 734.867 qps  |\n  | - [distill](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fastvit/fastvit_t12_distill.h5)      | 7.55M  | 1.34G | 256   | 80.3     | 734.867 qps  |\n  | - deploy=True | 7.50M  | 1.33G | 256   | 79.3     | 956.332 qps  |\n  | [FastViT_S12](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fastvit/fastvit_s12_imagenet.h5)   | 9.47M  | 1.74G | 256   | 79.9     | 666.669 qps  |\n  | - [distill](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fastvit/fastvit_s12_distill.h5)      | 9.47M  | 1.74G | 256   | 81.1     | 666.669 qps  |\n  | - deploy=True | 9.42M  | 1.74G | 256   | 79.9     | 881.429 qps  |\n  | [FastViT_SA12](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fastvit/fastvit_sa12_imagenet.h5) | 11.58M | 1.88G | 256   | 80.9     | 656.95 qps   |\n  | - [distill](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fastvit/fastvit_sa12_distill.h5)     | 11.58M | 1.88G | 256   | 81.9     | 656.95 qps   |\n  | - deploy=True | 11.54M | 1.88G | 256   | 80.9     | 833.011 qps  |\n  | [FastViT_SA24](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fastvit/fastvit_sa24_imagenet.h5) | 21.55M | 3.66G | 256   | 82.7     | 371.84 qps   |\n  | - [distill](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fastvit/fastvit_sa24_distill.h5)     | 21.55M | 3.66G | 256   | 83.4     | 371.84 qps   |\n  | - deploy=True | 21.49M | 3.66G | 256   | 82.7     | 444.055 qps  |\n  | [FastViT_SA36](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fastvit/fastvit_sa36_imagenet.h5) | 31.53M | 5.44G | 256   | 83.6     | 267.986 qps  |\n  | - [distill](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fastvit/fastvit_sa36_distill.h5)     | 31.53M | 5.44G | 256   | 84.2     | 267.986 qps  |\n  | - deploy=True | 31.44M | 5.43G | 256   | 83.6     | 325.967 qps  |\n  | [FastViT_MA36](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fastvit/fastvit_ma36_imagenet.h5) | 44.07M | 7.64G | 256   | 83.9     | 211.928 qps  |\n  | - [distill](https://github.com/leondgarse/keras_cv_attention_models/releases/download/fastvit/fastvit_ma36_distill.h5)     | 44.07M | 7.64G | 256   | 84.6     | 211.928 qps  |\n  | - deploy=True | 43.96M | 7.63G | 256   | 83.9     | 274.559 qps  |\n## FBNetV3\n  - [Keras FBNetV3](keras_cv_attention_models/mobilenetv3_family#fbnetv3) includes implementation of [PDF 2006.02049 FBNetV3: Joint Architecture-Recipe Search using Predictor Pretraining](https://arxiv.org/pdf/2006.02049.pdf).\n\n  | Model    | Params | FLOPs    | Input | Top1 Acc | T4 Inference |\n  | -------- | ------ | -------- | ----- | -------- | ------------ |\n  | [FBNetV3B](https://github.com/leondgarse/keras_cv_attention_models/releases/download/mobilenetv3_family/fbnetv3_b_imagenet.h5) | 5.57M  | 539.82M  | 256   | 79.15    | 713.882 qps  |\n  | [FBNetV3D](https://github.com/leondgarse/keras_cv_attention_models/releases/download/mobilenetv3_family/fbnetv3_d_imagenet.h5) | 10.31M | 665.02M  | 256   | 79.68    | 635.963 qps  |\n  | [FBNetV3G](https://github.com/leondgarse/keras_cv_attention_models/releases/download/mobilenetv3_family/fbnetv3_g_imagenet.h5) | 16.62M | 1379.30M | 256   | 82.05    | 478.835 qps  |\n## FlexiViT\n  - [Keras FlexiViT](keras_cv_attention_models/beit) includes models from [PDF 2212.08013 FlexiViT: One Model for All Patch Sizes](https://arxiv.org/pdf/2212.08013.pdf).\n\n  | Model         | Params  | FLOPs  | Input | Top1 Acc | T4 Inference |\n  | ------------- | ------- | ------ | ----- | -------- | ------------ |\n  | [FlexiViTSmall](https://github.com/leondgarse/keras_cv_attention_models/releases/download/beit/flexivit_small_240_imagenet.h5) | 22.06M  | 5.36G  | 240   | 82.53    | 744.578 qps  |\n  | [FlexiViTBase](https://github.com/leondgarse/keras_cv_attention_models/releases/download/beit/flexivit_base_240_imagenet.h5)  | 86.59M  | 20.33G | 240   | 84.66    | 301.948 qps  |\n  | [FlexiViTLarge](https://github.com/leondgarse/keras_cv_attention_models/releases/download/beit/flexivit_large_240_imagenet.h5) | 304.47M | 71.09G | 240   | 85.64    | 105.187 qps  |\n## GCViT\n  - [Keras GCViT](keras_cv_attention_models/gcvit) includes implementation of [PDF 2206.09959 Global Context Vision Transformers](https://arxiv.org/pdf/2206.09959.pdf).\n\n  | Model           | Params | FLOPs  | Input | Top1 Acc | Download |\n  | --------------- | ------ | ------ | ----- | -------- | -------- |\n  | [GCViT_XXTiny](https://github.com/leondgarse/keras_cv_attention_models/releases/download/gcvit/gcvit_xx_tiny_224_imagenet.h5)    | 12.0M  | 2.15G  | 224   | 79.9     | 337.7 qps   |\n  | [GCViT_XTiny](https://github.com/leondgarse/keras_cv_attention_models/releases/download/gcvit/gcvit_x_tiny_224_imagenet.h5)     | 20.0M  | 2.96G  | 224   | 82.0     | 255.625 qps   |\n  | [GCViT_Tiny](https://github.com/leondgarse/keras_cv_attention_models/releases/download/gcvit/gcvit_tiny_224_imagenet.h5)      | 28.2M  | 4.83G  | 224   | 83.5     | 174.553 qps   |\n  | [GCViT_Tiny2](https://github.com/leondgarse/keras_cv_attention_models/releases/download/gcvit/gcvit_tiny2_224_imagenet.h5)     | 34.5M  | 6.28G  | 224   | 83.7     |  |\n  | [GCViT_Small](https://github.com/leondgarse/keras_cv_attention_models/releases/download/gcvit/gcvit_small_224_imagenet.h5)     | 51.1M  | 8.63G  | 224   | 84.3     | 131.577 qps   |\n  | [GCViT_Small2](https://github.com/leondgarse/keras_cv_attention_models/releases/download/gcvit/gcvit_small2_224_imagenet.h5)    | 68.6M  | 11.7G  | 224   | 84.8     |  |\n  | [GCViT_Base](https://github.com/leondgarse/keras_cv_attention_models/releases/do","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fleondgarse%2Fkeras_cv_attention_models","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fleondgarse%2Fkeras_cv_attention_models","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fleondgarse%2Fkeras_cv_attention_models/lists"}