I recently spent a day chasing a memory problem in a Go service. After every deploy, its memory would creep up, slowly, over days. The autoscaler added pods to keep pace. Then the next deploy reset everything, and the climb started over. I love this kind of problem. It’s detective work: gather clues, probe with tools like pprof, work out why it happened, then make sure it can’t happen again. This case had a Sherlock Holmes twist, the dog that didn’t bark in the night. The most important clue was something that never happened. Nothing leaked.
It didn’t always push the service to its pod ceiling. Some cycles it barely moved. But the shape never changed: a slow ramp, a cliff at the deploy, another ramp. Because a deploy always “fixed” it, nobody looked very hard. That pattern hid the problem for years.
Every byte was reachable and held exactly where the code asked for it to be held. The code just asked for far too much.
A sawtooth is easy to live with#
A slow climb that resets on deploy never pages anyone. It doesn’t OOM. The dashboards look fine most of the time. If you ship often, pods never live long enough to get into trouble. If you ship less often, the autoscaler absorbs it, and the cost shows up on the bill instead of in an incident channel.
That’s what makes it dangerous. A crash gets a ticket. A gentle slope that resets itself gets a shrug. Without anyone deciding it, your deploy schedule becomes your memory limit. Ship less often, during a holiday freeze or a quiet quarter, and the bug that was always there suddenly shows up.
Where I looked first, and why it was wrong#
I took a few wrong turns. Each one is common enough to be worth naming.
The latest release. Whatever shipped last gets blamed when someone finally notices. The quick check is to compare pods of the same age across several releases. If they grow at the same rate before and after, the release didn’t cause it.
The biggest box in the heap profile. A heap profile tells you where memory was allocated, not who is keeping it alive. A decoder showing up large doesn’t make the decoder the problem. Ask what the decoded thing was handed to. And a profile from a laptop, where the only traffic is you, can’t show anything that grows with the number of customers.
The busiest traffic. Volume is not retention. A handler that allocates a lot per request and keeps none of it costs garbage collector CPU, not retained memory. What matters is what survives the request.
What did track memory was how many different customers had touched certain endpoints on a given pod, not how many requests they sent. That pointed straight at a per-customer cache.
What the cache was doing#
The code was ordinary. On a cache miss, it fetched the customer’s data from another service, decoded the protobuf response, and stored the entire decoded message in an in-process LRU keyed by customer. The handlers then pulled a couple of small lookup maps out of it and ignored everything else.
So every customer that touched those endpoints left a large tree of structs, maps, and strings sitting on the pod. Most of it was fields nobody ever read.
The fix was one idea: cache what you use, not what you fetched. Build the small maps once, cache those, and let the full message go. Same keys, same cache, same TTL, same call sites. Per-customer memory dropped by orders of magnitude, and the sawtooth flattened out.
Why the TTL didn’t save us#
The cache had a TTL. So why did entries stick around?
Because Go’s garbage collector doesn’t know what a TTL is. It frees an object once nothing reachable points to it. An expired entry is still referenced by the cache’s internal map, and the cache lives as long as the process. Reachable means alive. “Expired” is just a timestamp in a struct that nobody is reading.
In a simple LRU, expiry is checked on read. When you ask for a key, the cache looks at the timestamp and drops the entry if it’s stale. If nobody asks for that key again, nothing ever checks. A customer whose users don’t come back to that particular pod leaves their entry in place until the size cap pushes it out. In that kind of cache, the cap is the real memory limit, not the TTL.
Some libraries are sneakier. Their read path reports a stale entry as a miss but leaves it in memory until a background sweep gets to it. So even a read doesn’t free anything.
Here is a small program that shows it. It’s a bare-bones LRU with a TTL, not anyone’s production code. It forces a full garbage collection before each measurement, so the numbers only count memory that is still reachable:
package main
import (
"container/list"
"fmt"
"runtime"
"time"
)
// item is one cached value plus the moment it stops being valid.
type item struct {
key string
blob []byte
staleAt time.Time
}
// TinyLRU is a bare-bones LRU with a TTL. Like many simple caches,
// it only notices expiry when someone reads the key.
type TinyLRU struct {
capacity int
ttl time.Duration
order *list.List // front = most recently used
index map[string]*list.Element // key -> position in order
}
func NewTinyLRU(capacity int, ttl time.Duration) *TinyLRU {
return &TinyLRU{capacity: capacity, ttl: ttl, order: list.New(), index: make(map[string]*list.Element)}
}
func (c *TinyLRU) Put(key string, blob []byte) {
if el, found := c.index[key]; found {
it := el.Value.(*item)
it.blob, it.staleAt = blob, time.Now().Add(c.ttl)
c.order.MoveToFront(el)
return
}
c.index[key] = c.order.PushFront(&item{key: key, blob: blob, staleAt: time.Now().Add(c.ttl)})
if c.order.Len() > c.capacity { // the size cap: the only other way out
c.drop(c.order.Back())
}
}
func (c *TinyLRU) Lookup(key string) ([]byte, bool) {
el, found := c.index[key]
if !found {
return nil, false
}
it := el.Value.(*item)
if time.Now().After(it.staleAt) { // expiry is only checked here, on read
c.drop(el)
return nil, false
}
c.order.MoveToFront(el)
return it.blob, true
}
func (c *TinyLRU) drop(el *list.Element) {
c.order.Remove(el)
delete(c.index, el.Value.(*item).key)
}
func (c *TinyLRU) Size() int { return c.order.Len() }
// liveMB forces a full GC, then reports heap still in use.
// Anything still counted here is reachable.
func liveMB() float64 {
runtime.GC()
var ms runtime.MemStats
runtime.ReadMemStats(&ms)
return float64(ms.HeapAlloc) / (1 << 20)
}
func main() {
c := NewTinyLRU(100, time.Second)
fmt.Printf("empty cache heap=%6.1f MB\n", liveMB())
for i := 0; i < 50; i++ {
c.Put(fmt.Sprintf("customer-%d", i), make([]byte, 2<<20)) // a 2 MB "decoded response"
}
fmt.Printf("50 customers cached heap=%6.1f MB size=%d\n", liveMB(), c.Size())
time.Sleep(2 * time.Second)
fmt.Printf("all TTLs expired heap=%6.1f MB size=%d\n", liveMB(), c.Size())
_, ok := c.Lookup("customer-3")
fmt.Printf("read one key (hit=%v) heap=%6.1f MB size=%d\n", ok, liveMB(), c.Size())
for i := 50; i < 200; i++ {
c.Put(fmt.Sprintf("customer-%d", i), make([]byte, 2<<20))
}
fmt.Printf("200 customers seen heap=%6.1f MB size=%d\n", liveMB(), c.Size())
c = nil
fmt.Printf("cache dropped heap=%6.1f MB\n", liveMB())
}Output:
empty cache heap= 0.3 MB
50 customers cached heap= 100.3 MB size=50
all TTLs expired heap= 100.3 MB size=50
read one key (hit=false) heap= 98.3 MB size=49
200 customers seen heap= 200.4 MB size=100
cache dropped heap= 0.3 MBLook at the third line. Every entry has expired, the collector just ran, and the heap didn’t move. The fourth line is the entire expiry mechanism: reading one stale key frees that one entry. The fifth line is the cap doing the only other cleanup. The last line shows the collector works fine once the cache itself is gone.
Many TTL cache libraries add a background sweep to close that gap. Read the docs for yours, because the defaults vary. The expirable LRU in hashicorp/golang-lru runs a goroutine that deletes expired entries on a timer, as long as you set a TTL. go-cache only runs its janitor if you pass a cleanup interval, and it has no size cap at all. ttlcache leaves automatic cleanup off until you call Start(), and by default every hit pushes the expiry further out. A hand-rolled LRU usually has no sweep at all.
Even a working sweep only limits how long you hold an entry. It doesn’t make the entry any smaller. And a cache that extends the TTL on every hit keeps popular keys forever.
The garbage collector adds one more multiplier on top. GOGC sets how far the heap may grow past the live data before the next collection, which is double the live data at the default setting. The Go GC guide walks through the math, and GOMEMLIMIT can cap it. Whatever the cache holds, the container’s memory metric reads higher than the profile’s live heap. The autoscaler sees the container number.
Size isn’t only a memory problem, either. Everything reachable gets scanned on every collection. Discord hit the CPU side of this: their Go service’s big LRU cache made every garbage collection scan the whole thing, and that caused latency spikes.
And the collector never had to free any of it. When a pod restarts, the process exits and the operating system takes everything back at once. In a long-running service, that happens once per pod: on deploy. Which is why every deploy looked like a fix.
What I’d take away#
The lesson I keep coming back to is the smallest one: cache what you use, not what you fetched. If a handler derives something small from a big response, cache the small thing. Then do the multiplication. Entry size times the cap is a memory budget, and garbage collector headroom goes on top.
Next, know what your TTL actually does. Find out whether expiry happens only on read, whether anything sweeps, and whether that sweep is turned on. A long-lived in-process cache holds whatever it references for as long as the process runs. A TTL alone doesn’t change that.
Last, treat a memory slope that resets on deploy as something to investigate. It means your release schedule is acting as a memory limit nobody chose.
Nothing leaked. The cache kept exactly what it was told to keep. So how many caches in your services are storing the whole response because that was easier than picking out the two fields they actually use?
Reply by Email