{"id":13647893,"url":"https://github.com/jvehent/haproxy-aws","last_synced_at":"2025-04-22T06:32:15.766Z","repository":{"id":14734831,"uuid":"17455690","full_name":"jvehent/haproxy-aws","owner":"jvehent","description":"Documentation on building a HTTPS stack in AWS with HAProxy","archived":false,"fork":false,"pushed_at":"2015-11-18T02:43:30.000Z","size":97,"stargazers_count":280,"open_issues_count":0,"forks_count":64,"subscribers_count":30,"default_branch":"master","last_synced_at":"2024-11-09T22:36:08.540Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":"Shell","has_issues":false,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/jvehent.png","metadata":{"files":{"readme":"README.rst","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null}},"created_at":"2014-03-05T21:20:25.000Z","updated_at":"2024-02-22T19:45:57.000Z","dependencies_parsed_at":"2022-08-31T16:32:02.570Z","dependency_job_id":null,"html_url":"https://github.com/jvehent/haproxy-aws","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/jvehent%2Fhaproxy-aws","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/jvehent%2Fhaproxy-aws/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/jvehent%2Fhaproxy-aws/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/jvehent%2Fhaproxy-aws/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/jvehent","download_url":"https://codeload.github.com/jvehent/haproxy-aws/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":250183340,"owners_count":21388701,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-08-02T01:03:49.422Z","updated_at":"2025-04-22T06:32:15.453Z","avatar_url":"https://github.com/jvehent.png","language":"Shell","funding_links":[],"categories":["Open Source Repos","Shell"],"sub_categories":["Miscellaneous Repos"],"readme":"=========================================\nGuidelines for HAProxy termination in AWS\n=========================================\n\n.. title:: Guidelines for HAProxy termination in AWS\n.. |gentime| date:: %F %H:%M %Z\n.. role:: lvl_low\n\t:class: st_gray\n.. role:: lvl_medium\n\t:class: st_blue\n.. role:: lvl_high\n\t:class: st_yellow\n.. role:: lvl_max\n\t:class: st_red\n.. role:: ready\n\t:class: st_green\n.. role:: not_ready\n\t:class: st_red\n.. |br| raw:: html\n\n\t\u003cbr /\u003e\n\n.. sidebar:: Document status\n\n           +-----------------------+------------------------------------------------------+\n           |:not_ready:`NOT READY` | $Revision:        $ @ |gentime|                      |\n           +===============+=======+=============+=======================+================+\n           |**Author**     |Julien Vehent        |**Review**             | CloudOps       |\n           +---------------+---------------------+-----------------------+----------------+\n\n    .. sectnum::\n\n    .. contents:: **Table of contents**\n               :depth: 2\n\nSummary \u0026 Scope\n---------------\n\nThis document explains how HAProxy and Elastic Load Balancer can be used in\nAmazon Web Services to provide performant and secure termination of traffic\nto an API service. The goal is to provide the following features:\n\n- **DDoS Protection**: we use HAProxy to mitigate low to medium DDoS attacks, with\n  sane limits and custom blacklist.\n\n- **Application firewall**:  we perform a first level of filtering in HAProxy, that\n  protects NodeJS against all sorts of attack, known and to come. This will be done\n  by inserting a set of regexes in HAProxy ACLs, that get updated when the\n  application routes are updated. Note that managing these ACLs will not impact\n  uptime, or require redeployment.\n\n- **SSL/TLS**: ELBs support the PROXY protocol, and so does HAProxy, which allows us\n  to proxy the tcp connection to HAProxy. It gives us better TLS, backed by\n  OpenSSL, at the cost of managing the TLS keys on the HAProxy instances.\n\n- **Logging**: ELBs have limited support for logging. HAProxy, however, has excellent\n  logging for TCP, SSL and HTTPS. We leverage the flexibility of HAProxy's logging\n  to improve our DDoS detection capabilities. We also want to uniquely identify\n  requests in HAProxy and NodeJS, and correlate events, using a `unique-id`.\n\nArchitecture\n------------\n\nBelow is our target setup:\n\n.. image:: haproxy-aws-arch-diagram.png\n   :alt: architecture diagram\n\nPROXY protocol between ELB and HAProxy\n--------------------------------------\n\nThis configuration uses an Elastic Load Balancer in TCP mode, with PROXY\nprotocol enabled. The PROXY protocol adds a string at the beginning of the TCP\npayload that is passed to the backend. This string contains the IP of the client\nthat connected to the ELB, which allows HAProxy to feed its internal state with\nthis information, and act as if it had a direct TCP connection to the client.\n\nFor more information on the PROXY protocol, see\nhttp://haproxy.1wt.eu/download/1.5/doc/proxy-protocol.txt\n\nFirst, we need to create an ELB, and enable a TCP listener on port 443 that\nsupports the PROXY protocol. The ELB will not decipher the SSL, but instead pass\nthe entire TCP payload down to Haproxy.\n\nELB Configuration\n~~~~~~~~~~~~~~~~~\nPROXY protocol support must be enabled on the ELB.\n\n.. code:: bash\n\n    $ ./elb-describe-lb-policy-types -I AKIA... -S Ww1... --region us-east-1\n    POLICY_TYPE  ProxyProtocolPolicyType    Policy that controls whether to include the\n                                            IP address and port of the originating request\n                                            for TCP messages. This policy operates on\n                                            TCP/SSL listeners only\n\nThe policy name we want to enable is `ProxyProtocolPolicyType`. We need the load\nbalancer name for that, and the following command:\n\n.. code:: bash\n\n\t$ ./elb-create-lb-policy elb123-testproxyprotocol \\\n\t--policy-name EnableProxyProtocol \\\n\t--policy-type ProxyProtocolPolicyType \\\n\t--attribute \"name=ProxyProtocol, value=true\" \\\n\t-I AKIA... -S Ww1... --region us-east-1\n\n\tOK-Creating LoadBalancer Policy\n\n\n\t$ ./elb-set-lb-policies-for-backend-server elb123-testproxyprotocol \\\n\t--policy-names EnableProxyProtocol \\\n\t--instance-port 443 \\\n\t-I AKIA... -S Ww1... --region us-east-1\n\n\tOK-Setting Policies\n\nNow configure a listener on TCP/443 on that ELB, that points to TCP/443 on the\nHAProxy instance. On the instance side, make sure that your security group\naccepts traffic from the ELB security group on port 443.\n\nHAProxy frontend\n~~~~~~~~~~~~~~~~\n\nThe HAProxy frontend listens on port 443 with a SSL configuration, as follow:\n\n.. code::\n\n\tfrontend https\n\t\tbind 0.0.0.0:443 accept-proxy ssl ......\n\nNote the `accept-proxy` parameter of the bind command. This option tells HAProxy\nthat whatever sits in front of it will append the PROXY header to TCP payloads.\n\nSSL/TLS Configuration\n~~~~~~~~~~~~~~~~~~~~~\n\nHAProxy takes a SSL configuration on the `bind` line directly. The configuration\nrequires a set of certificates and private key, and a ciphersuite.\n\n.. code::\n\n\tbind 0.0.0.0:443 accept-proxy ssl crt /etc/haproxy/bundle.pem ciphers ECDHE-RSA-AES128-GCM-SHA256:ECDHE-ECDSA-AES128-GCM-SHA256:ECDHE-RSA-AES256-GCM-SHA384:ECDHE-ECDSA-AES256-GCM-SHA384:DHE-RSA-AES128-GCM-SHA256:DHE-DSS-AES128-GCM-SHA256:kEDH+AESGCM:ECDHE-RSA-AES128-SHA256:ECDHE-ECDSA-AES128-SHA256:ECDHE-RSA-AES128-SHA:ECDHE-ECDSA-AES128-SHA:ECDHE-RSA-AES256-SHA384:ECDHE-ECDSA-AES256-SHA384:ECDHE-RSA-AES256-SHA:ECDHE-ECDSA-AES256-SHA:DHE-RSA-AES128-SHA256:DHE-RSA-AES128-SHA:DHE-DSS-AES128-SHA256:DHE-RSA-AES256-SHA256:DHE-DSS-AES256-SHA:DHE-RSA-AES256-SHA:AES128-GCM-SHA256:AES256-GCM-SHA384:ECDHE-RSA-RC4-SHA:ECDHE-ECDSA-RC4-SHA:AES128:AES256:RC4-SHA:HIGH:!aNULL:!eNULL:!EXPORT:!DES:!3DES:!MD5:!PSK\n\nUnlike most servers (Apache, Nginx, ...), HAProxy takes certificates and keys\ninto a single file, here named `bundle.pem`. In this file are concatenated the\nserver private key, the server public certificate, the CA intermediate\ncertificate (if any) and a DH parameter (if any). For more information on DH\nparameters, see https://wiki.mozilla.org/Security/Server_Side_TLS .\n\nIn the sample below, components of `bundle.pem` are concatenated as follow:\n\n- client certificate signed by CA XYZ\n- client private key\n- public DH parameter (2048 bits)\n- intermediate certificate of CA XYZ\n\n.. code::\n\n\t-----BEGIN CERTIFICATE-----\n\tMIIGYjCCBUqgAwIBAgIDDD5PMA0GCSqGSIb3DQEBBQUAMIGMMQswCQYDVQQGEwJJ\n\t...\n\tej2w/mPv\n\t-----END CERTIFICATE-----\n\t-----BEGIN RSA PRIVATE KEY-----\n\tMIIEpAIBAAKCAQEAvJQqCjE4I63S3kR9KV0EG9e/lX/bZxa/2QVvZGi9/Suj65nD\n\t...\n\tRMSEpg+wuIVnKUi6KThiMKyXfZaTX7BDuR/ezE/JHs1TN5Hkw43TCQ==\n\t-----END RSA PRIVATE KEY-----\n\t-----BEGIN DH PARAMETERS-----\n\tMIICCAKCAgEA51RNlgY6j9MhmDURTpzydlJOsjk/TpU1BiY028SXAppuKJeFcx9S\n\t...\n\tHgHeuQQRjuv+h+Wf4dBe2f/fU5w9Osvq299vtcCjvQ7EtZTKT8RfvIMCAQI=\n\t-----END DH PARAMETERS-----\n\t-----BEGIN CERTIFICATE-----\n\tMIIGNDCCBBygAwIBAgIBGDANBgkqhkiG9w0BAQUFADB9MQswCQYDVQQGEwJJTDEW\n\t...\n\t0q6Dp6jOW6c=\n\t-----END CERTIFICATE-----\n\nThe rest of the `bind` line is a ciphersuite, taken from\nhttps://wiki.mozilla.org/Security/Server_Side_TLS .\n\nWe can verify the configuration using `cipherscan`. Below is the expected output\nfor our configuration:\n\n.. code:: bash\n\n\t$ ./cipherscan haproxytest1234.elb.amazonaws.com\n\t.........................\n\tprio  ciphersuite                  protocols                    pfs_keysize\n\t1     ECDHE-RSA-AES128-GCM-SHA256  SSLv3,TLSv1,TLSv1.1,TLSv1.2  ECDH,P-256,256bits\n\t2     ECDHE-RSA-AES256-GCM-SHA384  SSLv3,TLSv1,TLSv1.1,TLSv1.2  ECDH,P-256,256bits\n\t3     DHE-RSA-AES128-GCM-SHA256    SSLv3,TLSv1,TLSv1.1,TLSv1.2  DH,2048bits\n\t4     DHE-RSA-AES256-GCM-SHA384    SSLv3,TLSv1,TLSv1.1,TLSv1.2  DH,2048bits\n\t5     ECDHE-RSA-AES128-SHA256      SSLv3,TLSv1,TLSv1.1,TLSv1.2  ECDH,P-256,256bits\n\t6     ECDHE-RSA-AES128-SHA         SSLv3,TLSv1,TLSv1.1,TLSv1.2  ECDH,P-256,256bits\n\t7     ECDHE-RSA-AES256-SHA384      SSLv3,TLSv1,TLSv1.1,TLSv1.2  ECDH,P-256,256bits\n\t8     ECDHE-RSA-AES256-SHA         SSLv3,TLSv1,TLSv1.1,TLSv1.2  ECDH,P-256,256bits\n\t9     DHE-RSA-AES128-SHA256        SSLv3,TLSv1,TLSv1.1,TLSv1.2  DH,2048bits\n\t10    DHE-RSA-AES128-SHA           SSLv3,TLSv1,TLSv1.1,TLSv1.2  DH,2048bits\n\t11    DHE-RSA-AES256-SHA256        SSLv3,TLSv1,TLSv1.1,TLSv1.2  DH,2048bits\n\t12    DHE-RSA-AES256-SHA           SSLv3,TLSv1,TLSv1.1,TLSv1.2  DH,2048bits\n\t13    AES128-GCM-SHA256            SSLv3,TLSv1,TLSv1.1,TLSv1.2\n\t14    AES256-GCM-SHA384            SSLv3,TLSv1,TLSv1.1,TLSv1.2\n\t15    ECDHE-RSA-RC4-SHA            SSLv3,TLSv1,TLSv1.1,TLSv1.2  ECDH,P-256,256bits\n\t16    AES128-SHA256                SSLv3,TLSv1,TLSv1.1,TLSv1.2\n\t17    AES128-SHA                   SSLv3,TLSv1,TLSv1.1,TLSv1.2\n\t18    AES256-SHA256                SSLv3,TLSv1,TLSv1.1,TLSv1.2\n\t19    AES256-SHA                   SSLv3,TLSv1,TLSv1.1,TLSv1.2\n\t20    RC4-SHA                      SSLv3,TLSv1,TLSv1.1,TLSv1.2\n\t21    DHE-RSA-CAMELLIA256-SHA      SSLv3,TLSv1,TLSv1.1,TLSv1.2  DH,2048bits\n\t22    CAMELLIA256-SHA              SSLv3,TLSv1,TLSv1.1,TLSv1.2\n\t23    DHE-RSA-CAMELLIA128-SHA      SSLv3,TLSv1,TLSv1.1,TLSv1.2  DH,2048bits\n\t24    CAMELLIA128-SHA              SSLv3,TLSv1,TLSv1.1,TLSv1.2\n\nHealthchecks between ELB and HAProxy\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n\nAs of writing of this document, it appears that ELBs do not use the proxy\nprotocol when running healthchecks against an instance. As a result, these\nhealthchecks cannot be handled by the `https frontend`, because HAProxy will\nfail when looking for a PROXY header that isn't there.\n\nThe workaround is to create a secondary `frontend` in HAProxy that is entirely\ndedicated to answering healthchecks from the ELB.\n\nThe configuration below uses the `monitor` option to check the health of the\nnodejs backend. If more than one server is alive in that backend, then our\n`health` frontend will return `200 OK`. If no server is alive, a `503` will be\nreturned. All the ELB has to do is to query the URL at\nhttp://haproxy_host:34180/haproxy_status . To reduce the overhead, we also\ndisable SSL on the health frontend.\n\n.. code::\n\n\t# frontend used to return health status without requiring SSL\n\tfrontend health\n\t\tbind 0.0.0.0:34180\t# 34180 means EALTH ;)\n\t\t# create a status URI in /haproxy_status that will return\n\t\t# a 200 is backend is healthy, and 503 if it isn't. This\n\t\t# URI is queried by the ELB.\n\t\tacl backend_dead nbsrv(nodejs) lt 1\n\t\tmonitor-uri /haproxy_status\n\t\tmonitor fail if backend_dead\n\n*(note: we could also use ACLs in HAProxy to only expect the PROXY header on\ncertain source IPs, but the approach of a dedicated health frontend seems\ncleaner)*\n\nELB Logging\n-----------\nTODO\n\nHAProxy Logging\n---------------\n\nHAProxy supports custom log format, which we want here, as opposed to default\nlog format, in order to capture TCP, SSL and HTTP information on a single line.\n\nFor our logging, we want the following:\n\n1. TCP/IP logs first, such that these are always present, even if HAProxy cuts\n   the connection before processing the SSL or HTTP traffic\n2. SSL information\n3. HTTP information\n\n.. code::\n\n\tlog-format [%pid]\\ [%Ts.%ms]\\ %ac/%fc/%bc/%bq/%sc/%sq/%rc\\ %Tq/%Tw/%Tc/%Tr/%Tt\\ %tsc\\ %ci:%cp\\ %fi:%fp\\ %si:%sp\\ %ft\\ %sslc\\ %sslv\\ %{+Q}r\\ %ST\\ %b:%s\\ \"%CC\"\\ \"%hr\"\\ \"%CS\"\\ \"%hs\"\\ req_size=%U\\ resp_size=%B\n\nThe format above will generate:\n\n.. code::\n\n\tMar 14 17:14:51 localhost haproxy[14887]: [14887] [1394817291.250] 10/5/2/0/3/0/0 48/0/0/624/672 ---- 1.10.2.10:35701 10.151.122.228:443 127.0.0.1:8000 logger - - \"GET /v1/ HTTP/1.0\" 404 fxa-nodejs:nodejs1 \"-\" \"{||ApacheBench/2.3|over-100-active-connections,over-100-connections-in-10-seconds,high-error-rate,high-request-rate,|47B4176E:8B75_0A977AE4:01BB_5323390B_31E0:3A27}\" \"-\" \"\" ireq_size=592 resp_size=787\n\nThe log-format contains very detailed information on the connection itself, but\nalso on the state of haproxy itself. Below is a description of the fields we\nused in our custom log format.\n\n* `%pid`: process ID of HAProxy\n* `%Ts.%ms`: unix timestamp + milliseconds\n* `%ac`: total number of concurrent connections\n* `%fc`: total number of concurrent connections on the frontend\n* `%bc`: total number of concurrent connections on the backend\n* `%bq`: queue size of the backend\n* `%sc`: total number of concurrent connections on the server\n* `%sq`: queue size of the server\n* `%rc`: connection retries to the server\n* `%Tq`: total time to get the client request (HTTP mode only)\n* `%Tw`: total time spent in the queues waiting for a connection slot\n* `%Tc`: total time to establish the TCP connection to the server\n* `%Tr`: server response time (HTTP mode only)\n* `%Tt`: total session duration time, between the moment the proxy accepted it\n  and the moment both ends were closed.\n* `%tsc`: termination state (see `8.5. Session state at disconnection`)\n* `%ci:%cp`: client IP and Port\n* `%fi:%fp`: frontend IP and Port\n* `%si:%sp`: server IP and Port\n* `%ft`: transport type of the frontend (with a ~ suffix for SSL)\n* `%sslc %sslv`: SSL cipher and version\n* `%{+Q}r`: HTTP request, between double quotes\n* `%ST`: HTTP status code\n* `%b:%s`: backend name and server name\n* `%CC`: captured request cookies\n* `%hr`: captured request headers\n* `%CS`: captured response cookies\n* `%hs`: captured response headers\n* `%U`: bytes read from the client (request size)\n* `%B`: bytes read from server to client (response size)\n\nFor more details on the available logging variables, see the HAProxy\nconfiguration, under `8.2.4. Custom log format`.\nhttp://haproxy.1wt.eu/download/1.5/doc/configuration.txt\n\nUnique request ID\n~~~~~~~~~~~~~~~~~\n\nTracking requests across multiple servers can be problematic, because the chain\nof events triggered by a request on the frontend are not tied to each other.\nHAProxy has a simple mechanism to insert a unique identifier to incoming\nrequests, in the form of an ID inserted in the request headers, and passed to\nthe backend server. This ID can then be logged by the backend server, and passed\non to the next step. In a largely distributed environment, the unique ID makes\ntracking requests propagation a lot easier.\n\nThe unique ID is declared on the HTTPS frontend as follow:\n\n.. code::\n\n\t# Insert a unique request identifier is the headers of the request\n\t# passed to the backend\n\tunique-id-format %{+X}o\\ %ci:%cp_%fi:%fp_%Ts_%rt:%pid\n\tunique-id-header X-Unique-ID\n\nThis will add an ID that is composed of hexadecimal variables, taken from the\nclient IP and port, frontend IP and port, timestamp, request counter and PID.\nAn example of generated ID is **485B7525:CB2F_0A977AE4:01BB_5319CB0C_000D:27C0**.\n\nThe Unique ID is added to the request headers passed to the backend in the\n`X-Unique-ID` header. We will also capture it in the logs, as a request header.\n\n ::\n\n\tGET / HTTP/1.1\n\tHost: backendserver123.example.net\n\tUser-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:25.0) Gecko/20100101 Firefox/25.0\n\tAccept: text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8\n\tAccept-Language: en-US,en;q=0.5\n\tAccept-Encoding: gzip, deflate\n\tDNT: 1\n\tCache-Control: max-age=0\n\tX-Unique-ID: 485B7525:CB70_0A977AE4:01BB_5319CD3F_0163:27C0\n\tX-Forwarded-For: 2.12.17.87\n\nCapturing headers and cookies\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n\nIn the log format, we defined fields for the request and response headers and\ncookies. But, by default, this fields will show empty in the logs. In order to\nlog headers and cookies, the `capture` parameters must be set in the\nfrontend.\n\nHere is how we can capture headers sent by the client in the HTTP request.\n\n.. code::\n\n\tcapture request header Referrer len 64\n    capture request header Content-Length len 10\n\tcapture request header User-Agent len 64\n\nCookies can be captures the same way:\n\n.. code::\n\n\tcapture cookie mycookie123=  len 32\n\nHAProxy will also add custom headers to the request, before passing it to the\nbackend. However, added headers don't get logged, because the addition happens\nafter the capture operation. To fix this issue, we are going to create a new\nfrontend dedicated to logging.\n\nLogging in a separate frontend\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n\nDuring processing of the request, we added custom headers, and we want these\nheaders to appear in the logs. One solution is to route all the request to a\nsecondary frontend that only does logging, and blocking or forwarding.\n\nClassic setup:\n\n ::\n\n                    {logging}\n     request        +--------------+       +---------------+\n    +--------------\u003e|frontend      |+-----\u003e|backend        |      +---------+\n                    |   fxa-https  |       |    fxa-nodejs |+----\u003e|         |\n                    +--------------+       +---------------+      | NodeJS  |\n                                                                  |         |\n                                                                  +---------+\n\nSetup with separate logging frontend:\n\n ::\n\n                    {no logging}\n     request        +--------------+       +---------------+\n    +--------------\u003e|frontend      |       |backend        |      +---------+\n                    |   fxa-https  |       |    fxa-nodejs |+----\u003e|         |\n                    +--------------+       +---------------+      | NodeJS  |\n                           +                     ^                |         |\n                           |                     |                +---------+\n                           |                     |\n                    +------v-------+       +-----+--------+\n                    |backend       |+-----\u003e|frontend      |\n                    |     logger   |       |   logger     |\n                    +--------------+       +--------------+\n                                             {logging}\n\n\nAt the end of the configuration of frontend `fxa-https`, instead of sending\nrequests to backend `fxa-nodejs`, we send them to backend `logger`.\n\n.. code::\n\n\tfrontend fxa-https\n\t\t...\n\t\t# Don't log here, log into logger frontend\n\t\tno log\n\t\tdefault_backend logger\n\nThen we declare a backend and a frontend for `logger`:\n\n.. code::\n\n\tbackend logger\n\t\tserver localhost localhost:55555 send-proxy\n\n\t# frontend use to log acl activity\n\tfrontend logger\n\t\tbind localhost:55555 accept-proxy\n\n\t\t...\n\n\t\tcapture request header Referrer len 64\n\t\tcapture request header Content-Length len 10\n\t\tcapture request header User-Agent len 64\n\t\tcapture request header X-Haproxy-ACL len 256\n\t\tcapture request header X-Unique-ID len 64\n\n\t\t# if previous ACL didn't pass and aren't whitelisted\n\t\tacl whitelisted req.fhdr(X-Haproxy-ACL) -m beg whitelisted,\n\t\tacl fail-validation req.fhdr(X-Haproxy-ACL) -m found\n\t\thttp-request deny if !whitelisted fail-validation\n\n\t\tdefault_backend fxa-nodejs\n\nNote the use of `send-proxy` and `accept-proxy` between the logger backend and\nfrontend, allowing to keep the information about the client IP.\n\n**Isn't this slow and inefficient?**\n\nWell, obviously, routing request through HAProxy twice isn't the most elegant\nway of proxying. But in practice, this approach adds minimal overhead. Linux and\nHAProxy support TCP splicing, which provides zero-copy transfer of data between\nTCP sockets. When HAProxy forward the request to the logger socket, there is, in\nfact, no transfer of data at the kernel level. Benchmark it, it's fast!\n\nRate limiting \u0026 DDoS protection\n-------------------------------\n\nOne of the particularity of operating an infrastructure in AWS, is that control\nover the network is very limited. Techniques such as BGP blackholing are not\navailable. And visibility over the layer 3 (IP) and 4 (TCP) is reduced. Building\nprotection against DDoS means that we need to block traffic further down the\nstack, which consumes more resources. This is the main motivation for using ELBs\nin TCP mode with the PROXY protocol: it gives HAProxy low-level access to the\nTCP connection, and visibility of the client IP before parsing HTTP headers\n(like you would traditionally do with X-Forwarded-For).\n\nELBs have limited resources, but simplify the management of public IPs in AWS.\nBy offloading the SSL \u0026 HTTP processing to HAProxy, we reduce the pressure on\nELB, while conserving the ability to manage the public endpoints through it.\n\nHAProxy maintains tons of detailed information on connections. One can use this\ninformation to accept, block or route connections. In the following section, we\nwill discuss the use of ACLs and stick-tables to block clients that do not\nrespect sane limits.\n\nAutomated rate limiting\n~~~~~~~~~~~~~~~~~~~~~~~\n\nThe configuration below enable counters to track connections in a table where\nthe key is the source IP of the client:\n\n.. code::\n\n\t# Define a table that will store IPs associated with counters\n\tstick-table type ip size 500k expire 30s store conn_cur,conn_rate(10s),http_req_rate(10s),http_err_rate(10s)\n\n\t# Enable tracking of src IP in the stick-table\n\ttcp-request content track-sc0 src\n\nLet's decompose this configuration. First, we define a `stick-table` that\nstores IP addresses as keys. We define a maximum size for this table\nof 500,000 IPs, and we tell HAProxy to expire the records after 30 seconds. If\nthe table gets filled, HAProxy will delete records following the LRU logic.\n\nThe `stick-table` will store a number of information associated with the IP\naddress:\n\n- `conn_cur` is a counter of the concurrent connection count for this IP.\n\n- `conn_rate(10s)` is a sliding window that counts new TCP connections over a 10\n  seconds period\n\n- `http_req_rate(10s)` is a sliding window that counts HTTP requests over a 10\n  seconds period\n\n- `http_err_rate(10s)` is a sliding window that counts HTTP errors triggered by\n  requests from that IP over a 10 seconds period\n\nBy default, the stick table declaration doesn't do anything, we need to send\ndata to it. This is what the `tcp-request content track-sc0 src` parameter does.\n\nNow that we have tracking in place, we can write ACLs that run tests against the\ncontent of the table. The examples below evaluate several of these counters\nagainst arbitrary limits. Tune these to your needs.\n\n.. code::\n\n\t# Reject the new connection if the client already has 100 opened\n\thttp-request add-header X-Haproxy-ACL %[req.fhdr(X-Haproxy-ACL,-1)]over-100-active-connections, if { src_conn_cur ge 100 }\n\n\t# Reject the new connection if the client has opened more than 100 connections in 10 seconds\n\thttp-request add-header X-Haproxy-ACL %[req.fhdr(X-Haproxy-ACL,-1)]over-100-connections-in-10-seconds, if { src_conn_rate ge 100 }\n\n\t# Reject the connection if the client has passed the HTTP error rate\n\thttp-request add-header X-Haproxy-ACL %[req.fhdr(X-Haproxy-ACL,-1)]high-error-rate, if { sc0_http_err_rate() gt 100 }\n\n\t# Reject the connection if the client has passed the HTTP request rate\n\thttp-request add-header X-Haproxy-ACL %[req.fhdr(X-Haproxy-ACL,-1)]high-request-rate, if { sc0_http_req_rate() gt 500 }\n\nHAProxy provides a lot of flexibility on what can be tracked in a `stick-table`.\nTake a look at section `7.3.2. Fetching samples at Layer 4` from the doc to get\na better idea.\n\nQuerying tables state in real time\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n\nTables are named after the name of the frontend or backend they live in. Our\nfrontend called `fxa-https` will have a table called `fxa-https`, that can be\nqueried through the stat socket:\n\n.. code::\n\n\t# echo \"show table fxa-https\" | socat unix:/var/lib/haproxy/stats -\n\t# table: fxa-https, type: ip, size:512000, used:1\n\t0x1aa3358: key=1.10.2.10 use=1 exp=29957 conn_rate(10000)=43 conn_cur=1 http_req_rate(10000)=42 http_err_rate(10000)=42\n\nThe line above shows a table entry for key `1.10.2.10`, which is a tracked IP\naddress. The other entries on the line show the status of various counters that\nwe defined in the configuration.\n\nBlacklists \u0026 Whitelists\n~~~~~~~~~~~~~~~~~~~~~~~\n\nBlacklist and whitelists are simple lists of IP addresses that are checked by\nHAProxy as early on as possible. Blacklist are checked at the beginning of the\nTCP connection, which allows for early connection drops, and also means that\nblacklisting an IP always takes precedence over any other rule, including the\nwhitelist.\n\nWhitelists are checked at the HTTP level, and allow to bypass ACLs and rate\nlimiting.\n\n.. code::\n\n\t# Blacklist: Deny access to some IPs before anything else is checked\n\ttcp-request content reject if { src -f /etc/haproxy/blacklist.lst }\n\n\t# Whitelist: Allow IPs to bypass the filters\n\thttp-request add-header X-Haproxy-ACL %[req.fhdr(X-Haproxy-ACL,-1)]whitelisted, if { src -f /etc/haproxy/whitelist.lst }\n\thttp-request allow if { src -f /etc/haproxy/whitelist.lst }\n\nList files can contain IP addresses or networks in CIDR format.\n\n.. code::\n\n\t10.0.0.0/8\n\t172.16.0.0/12\n\t192.168.0.0/16\n\t8.8.8.8\n\nList files are loaded into HAProxy at startup. If you add or remove IPs from a\nlist, make sure to perform a soft reload.\n\n.. code:: bash\n\n\thaproxy -f /etc/haproxy/haproxy.cfg -c \u0026\u0026 sudo haproxy -f /etc/haproxy/haproxy.cfg -sf $(pidof haproxy)\n\nProtect against slow clients (Slowloris attack)\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n\nSlowloris is an attack where a client very slowly sends requests to the server,\nforcing it to allocate resources to that client that are only not used. This\nattack is commonly used in DDoS, by clients that send their requests characters\nby characters. HAProxy can block these clients, by allocating a maximum amount\nof time a client can take to send a full request. This is done with the `timeout\nhttp-request` parameter.\n\n.. code::\n\n    # disconnect slow handshake clients early, protect from\n    # resources exhaustion attacks\n    timeout http-request 5s\n\nURL filtering with ACLs\n-----------------------\n\nHAProxy has the ability to inspect requests before passing them to the backend.\nThis is limited to query strings, and doesn't support inspecting the body of a\nPOST request. But we can already leverage this to filter out unwanted traffic.\n\nThe first thing we need, is a list of endpoints sorted by HTTP method. This can\nbe obtained from the web application directly. Note that some endpoints, such as\n`__heartbeat__` should be limited to HAProxy, and thus blocked from clients.\n\nFor now, let's ignore GET URL parameters, and only build a list of request\npaths, that we store in two files: one for GET requests, and one for POST\nrequests.\n\n`get_endpoints.lst`\n\n.. include :: get_endpoints.lst\n   :code: bash\n\n`post_endpoints.lst`\n\n.. include :: post_endpoints.lst\n   :code: bash\n\nIn the HAProxy configuration, we can build ACLs around these files. The `http-request deny`\nmethod takes a condition, as described in the Haproxy documentation, section\n`7.2. Using ACLs to form conditions`.\n\n.. code::\n\n\t# Requests validation using ACLs ---\n\tacl valid-get path -f /etc/haproxy/get_endpoints.lst\n\tacl valid-post path -f /etc/haproxy/post_endpoints.lst\n\n\t# block requests that don't match the predefined endpoints\n\thttp-request deny unless METH_GET valid-get or METH_POST valid-post\n\n`http-request deny` does the job, and return a 403 to the client. But if you want\nmore visibility on ACL activity, you may want to use a custom header as describe\nlater in this section.\n\nFiltering URL parameters on GET requests\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n\nWhile HAProxy supports regexes on URLs, writing regexes that can validate URL\nparameters is a path that leads to frustration and insanity. A much simpler\napproach consists of using the `url_param` ACL provided by HAProxy.\n\nFor example, take the NodeJS endpoint below:\n\n.. code:: javascript\n\n    {\n      method: 'GET',\n      path: '/verify_email',\n      config: {\n        validate: {\n          query: {\n            code: isA.string().max(32).regex(HEX_STRING).required(),\n            uid: isA.string().max(32).regex(HEX_STRING).required(),\n            service: isA.string().max(16).alphanum().optional(),\n            redirectTo: isA.string()\n              .max(512)\n              .regex(validators.domainRegex(redirectDomain))\n              .optional()\n          }\n        }\n      },\n      handler: function (request, reply) {\n        return reply().redirect(config.contentServer.url + request.raw.req.url)\n      }\n    },\n\nThis endpoints receives requests on `/verify_email` with the parameters `code`,\na 32 character hexadecimal, `uid`, a 32 character hexadecimal, `service`, a 16\ncharacter string, and `redirectTo`, a FQDN. However, only `code` and `uid` are\nrequired.\n\nIn the previous section, we validated that requests on `/verify_email` must use\nthe method GET. Now we are taking the validation one step further, and blocking\nrequests on this endpoint that do not match our prerequisite.\n\n.. code::\n\n\tacl endpoint-verify_email path /verify_email\n\tacl param-code urlp_reg(code) [0-9a-fA-F]{1,32}\n\tacl param-uid urlp_reg(uid) [0-9a-fA-F]{1,32}\n\thttp-request deny if endpoint-verify_email !param-code or endpoint-verify_email !param-uid\n\nThe follow request will be accepted, everything else will be rejected with a\nHTTP error 403.\n\n.. code::\n\n\thttps://haproxy_server/verify_email?code=d64f53326cec3a1af60166a929ca52bd\u0026uid=d64f53326cec3a1af60166a929c3d7b2131561792b4837377ed2e0cde3295df2\n\nUsing regexes to validate URL parameters is a powerful feature. Below is another\nexample that matches an email addresses using case-insensitive regex:\n\n.. code::\n\n\tacl endpoint-complete_reset_password path /complete_reset_password\n\tacl param-email urlp_reg(email) -i ^[A-Z0-9._%+-]+@[A-Z0-9.-]+\\.[A-Z]{2,4}$\n\tacl param-token urlp_reg(token) [0-9a-fA-F]{1,64}\n\thttp-request deny if endpoint-complete_reset_password !param-email or endpoint-complete_reset_password !param-token or endpoint-complete_reset_password !param-code\n\nNote that we didn't redefine `param-code` when we reused it in the `http-request deny`\ncommand. This is because ACL are defined globally for a frontend, and can\nbe reused multiple times.\n\nFiltering payloads on POST requests\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n\nPOST requests are harder to validate, because they do not follow a predefined\nformat, but also because the client could be sending the body over a long period\nof time, split over dozens of packets.\n\nHowever, in the case of an API that only handles small POST payloads, we can at\nleast verify the size of the payload sent by the client, and make sure that\nclients do not overload the backend with random data. This can be done using an\nACL on the content-length header of the request. The ACL below discard requests\nthat have a content-length larger than 5 kilo-bytes (which is already a lot of\ntext).\n\n.. code::\n\n\t# match content-length larger than 5kB\n\tacl request-too-big hdr_val(content-length) gt 5000\n\thttp-request deny if METH_POST request-too-big\n\nMarking instead of blocking\n~~~~~~~~~~~~~~~~~~~~~~~~~~~\n\nBlocking requests may be the preferred behavior in production, but only after a\ngrace period that allows you to build a traffic profile, and fine tune your\nconfiguration. Instead of using `http-request deny` statements in the ACLs, we\ncan insert a header with a description of the blocking decision. This header\nwill be logged, and can be analyzed to verify that no legitimate traffic would\nbe blocked.\n\nAs discussed in `Logging in a separate frontend`, HAProxy is unable to log\nrequest header that it has set itself. So make sure to log in a separate\nfrontend if you use this technique.\n\nThe configuration below uses a custom header `X-Haproxy-ACL`. If an ACL matches,\nthe header is set to the name of the ACL that matched. If several ACLs match,\neach ACL name is appended to the header, and separated by a comma.\n\nAt the end of the ACL evaluation, if this header is present in the request, we\nknow that the request should be blocked.\n\nIn the `fxa-https` frontend, we replace the `http-request deny` paramameters with the\nfollowing logic:\n\n.. code::\n\n\t# ~~~ Requests validation using ACLs ~~~\n\t# block requests that don't match the predefined endpoints\n\tacl valid-get path -f /etc/haproxy/get_endpoints.lst\n\tacl valid-post path -f /etc/haproxy/post_endpoints.lst\n\thttp-request add-header X-Haproxy-ACL %[req.fhdr(X-Haproxy-ACL,-1)]invalid-endpoint, unless METH_GET valid-get or METH_POST valid-post\n\n\t# block requests on verify_email that do not have the correct params\n\tacl endpoint-verify_email path /v1/verify_email\n\tacl param-code urlp_reg(code) [0-9a-fA-F]{1,32}\n\tacl param-uid urlp_reg(uid) [0-9a-fA-F]{1,32}\n\thttp-request add-header X-Haproxy-ACL %[req.fhdr(X-Haproxy-ACL,-1)]invalid-parameters, if endpoint-verify_email !param-code or endpoint-verify_email !param-uid\n\n\t# block requests on complete_reset_password that do not have the correct params\n\tacl endpoint-complete_reset_password path /v1/complete_reset_password\n\tacl param-email urlp_reg(email) -i ^[A-Z0-9._%+-]+@[A-Z0-9.-]+\\.[A-Z]{2,4}$\n\tacl param-token urlp_reg(token) [0-9a-fA-F]{1,64}\n\thttp-request add-header X-Haproxy-ACL %[req.fhdr(X-Haproxy-ACL,-1)]invalid-parameters, if endpoint-complete_reset_password !param-email or endpoint-complete_reset_password !param-token or endpoint-complete_reset_password !param-code\n\n\t# block content-length larger than 500kB\n\tacl request-too-big hdr_val(content-length) gt 5000\n\thttp-request add-header X-Haproxy-ACL %[req.fhdr(X-Haproxy-ACL,-1)]request-too-big, if METH_POST request-too-big\n\nNote the `%[req.fhdr(X-Haproxy-ACL,-1)]` parameter, that retrieves the value of\nthe latest occurence of the `X-Haproxy-ACL` header, so we can append to it and\nstore it again. However, this will create multiple headers if more than one ACL\nis matched, but that's OK because:\n- we can delete them before sending the request to the backend, using `reqdel`\n- the logging directive `capture request header` will only log the last occurence\n\n.. code::\n\n\tX-Haproxy-ACL: over-100-active-connections,\n\tX-Haproxy-ACL: over-100-active-connections,over-100-connections-in-10-seconds,\n\tX-Haproxy-ACL: over-100-active-connections,over-100-connections-in-10-seconds,high-error-rate,\n\tX-Haproxy-ACL: over-100-active-connections,over-100-connections-in-10-seconds,high-error-rate,high-request-rate,\n\nThen, in the logger frontend, we check the value of the header, and block if\nneeded.\n\n.. code::\n\n\t# frontend use to log acl activity\n\tfrontend logger\n\t\t...\n\t\t# if previous ACL didn't pass, and IP isn't whitelisted, block the request\n\t\tacl whitelisted req.fhdr(X-Haproxy-ACL) -m beg whitelisted,\n\t\tacl fail-validation req.fhdr(X-Haproxy-ACL) -m found\n\t\thttp-request deny if !whitelisted fail-validation\n\nHAProxy management\n------------------\n\nEnabling the stat socket\n~~~~~~~~~~~~~~~~~~~~~~~~\n\nCollecting statistics\n~~~~~~~~~~~~~~~~~~~~~\n\nAnalyzing errors\n~~~~~~~~~~~~~~~~\n\nParsing performance metrics from the logs\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n\nSoft reload\n~~~~~~~~~~~\nHAProxy supports soft configuration reload, that doesn't drop connections. To\nperform a soft reload, call haproxy with the following command:\n\n.. code:: bash\n\n\t$ sudo /opt/haproxy -f /etc/haproxy/haproxy.cfg -sf $(pidof haproxy)\n\nThe old process will be replaced with a new one, that uses a fresh\nconfiguration. The logs will show the reload:\n\n ::\n\n\tMar  6 12:59:41 localhost haproxy[7603]: Proxy https started.\n\tMar  6 12:59:41 localhost haproxy[7603]: Proxy app started.\n\tMar  6 12:59:41 localhost haproxy[5763]: Stopping frontend https in 0 ms.\n\tMar  6 12:59:41 localhost haproxy[5763]: Stopping backend app in 0 ms.\n\tMar  6 12:59:41 localhost haproxy[5763]: Proxy https stopped (FE: 29476 conns, BE: 0 conns).\n\tMar  6 12:59:41 localhost haproxy[5763]: Proxy app stopped (FE: 0 conns, BE: 1746 conns).\n\nFull HAProxy configuration\n--------------------------\n\n.. include:: haproxy.cfg\n   :code: bash\n\nBuilding process\n----------------\n\nStatic build\n~~~~~~~~~~~~\nThe script `build_static_haproxy.sh`_ builds haproxy with statically linked\nOpenSSL and PCRE support.\n\n.. _`build_static_haproxy.sh`: build_static_haproxy.sh\n\n.. include:: build_static_haproxy.sh\n\t:code: bash\n\nDynamic build\n~~~~~~~~~~~~~\nThe script `build_dynamic_haproxy.sh`_ does the same as above, but links to\nPCRE and OpenSSL dynamically.\n\n.. _`build_dynamic_haproxy.sh`: build_dynamic_haproxy.sh\n\n.. include:: build_dynamic_haproxy.sh\n   :code: bash\n\nRPM build\n~~~~~~~~~\nUsing the spec file in `haproxy.spec`_ and bash scripts in `build_rpm.sh`_,\nwe can build a RPM package using for the latest development version of HAProxy.\n\n.. _`haproxy.spec`: haproxy.spec\n\n.. _`build_rpm.sh`: build_rpm.sh\n\n`haproxy.spec`\n\n.. include:: haproxy.spec\n   :code: bash\n\n`build_rpm.sh`\n\n.. include:: build_rpm.sh\n   :code: bash\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fjvehent%2Fhaproxy-aws","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fjvehent%2Fhaproxy-aws","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fjvehent%2Fhaproxy-aws/lists"}