DB Write Protocols & Proxies

From my experience working at GreyOrange. Refactored my article a bit with help of GPT.

I was working on a component of a golang-service that intercepted InfluxDB writes (proxy), to enforce few rules like dropping unoptimised queries, emitting query metrics, mirroring writes to Kafka, etc. As a part of exploring the service, I learned about how different databases actually receive writes under the hood. The protocol a DB uses determines everything — I have discussed the same in this article at a high level.

There are two fundamentally different classes of DB write protocols.

Class 1: HTTP-Based Databases

InfluxDB, Elasticsearch, and CouchDB expose plain HTTP REST endpoints. A write is an HTTP request — nothing more.

Class 2: TCP Wire Protocol Databases

Postgres, MySQL, Redis, and MongoDB do not use HTTP. They define their own binary (or text) framing over a raw TCP socket. When your application writes, the driver opens a TCP connection and speaks the DB’s private protocol directly.

Why the Protocol Matters

DB Transport Body encoding Query/command in
InfluxDB HTTP Line Protocol (text) URL param ?q= or body
Elasticsearch HTTP JSON / NDJSON JSON body
CouchDB HTTP JSON URL path + JSON body
PostgreSQL TCP Binary frames Q message payload (text SQL)
MySQL TCP Binary frames COM_QUERY payload (text SQL)
Redis TCP RESP (text arrays) First element of RESP array
MongoDB TCP Binary frames + BSON BSON document body

This also determines how hard it is to debug:

If You Were to Build a Proxy in Front of These DBs

The Core Insight

HTTP-based DBs are just web services with domain-specific body formats. Any standard HTTP middleware — reverse proxies, API gateways, service meshes — can sit in front of them and inspect traffic without knowing anything about the DB.

TCP wire protocol DBs speak their own language. A proxy must implement or embed a parser for that binary protocol before it can read a single query. The upside: once you’ve parsed the framing, the query is just a string and the same filtering logic applies as for HTTP DBs.

The filtering logic itself — time bounds, rate limits, blocklists, cost estimation — is identical for both classes. The only difference is how much work you do to get the query into a readable form in the first place.