Web applications are increasingly targeted by attacks that exploit HTTP requests to bypass security controls. Traditional WAFs rely on rule‑based logic, which leads to high false positive rates and limited adaptability. Recent work has applied machine learning and word embeddings to improve anomaly detection. This paper builds a unified single‑class classification benchmark that compares three static embedding models: Word2Vec, FastText, and Doc2Vec. We introduce HEDA (HTTP Embedding‑Based Detection Architecture), which feeds the embedding vectors into a single‑class anomaly detector to evaluate each request. The entire pipeline operates unsupervised, training both the embedding models and detectors solely on benign traffic. Experiments span three heterogeneous datasets, including synthetic and real traffic. Findings show that the choice of embedding has a decisive impact on performance; FastText consistently yields the best results, achieving high detection rates while keeping false positives low.
Review