<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Troubleshooting on Thalamus</title><link>https://blog.thalamus.am/tag/troubleshooting/</link><description>Recent content in Troubleshooting on Thalamus</description><generator>Hugo -- gohugo.io</generator><language>en-US</language><copyright>Thalamus TM</copyright><lastBuildDate>Sun, 20 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://blog.thalamus.am/tag/troubleshooting/index.xml" rel="self" type="application/rss+xml"/><item><title>The 200 that nobody received</title><link>https://blog.thalamus.am/posts/the-200-that-nobody-received/</link><pubDate>Sun, 20 Sep 2026 00:00:00 +0000</pubDate><guid>https://blog.thalamus.am/posts/the-200-that-nobody-received/</guid><description>&lt;img src="https://blog.thalamus.am/posts/the-200-that-nobody-received/editorial-cover.png" alt="Featured image of post The 200 that nobody received" /&gt;&lt;p&gt;Clients of &lt;code&gt;api.thalamus.am&lt;/code&gt; were occasionally timing out. The gateway access log showed successful requests.&lt;/p&gt;
&lt;p&gt;The request had arrived. The upstream had answered. The log recorded a 200. There was no obvious pattern by endpoint or time of day, and the failures were rare enough to disappear into otherwise healthy dashboards.&lt;/p&gt;
&lt;p&gt;I spent too long looking for an application error. The useful question turned out to be: which machines did the connection depend on after the application had done its work?&lt;/p&gt;</description></item><item><title>df says full, du disagrees</title><link>https://blog.thalamus.am/posts/df-says-full-du-disagrees/</link><pubDate>Sat, 08 Aug 2026 00:00:00 +0000</pubDate><guid>https://blog.thalamus.am/posts/df-says-full-du-disagrees/</guid><description>&lt;img src="https://blog.thalamus.am/posts/df-says-full-du-disagrees/editorial-cover.png" alt="Featured image of post df says full, du disagrees" /&gt;&lt;p&gt;The third time the compactor volume filled up, I checked what was using the space before making it bigger.&lt;/p&gt;
&lt;p&gt;The alert was &lt;code&gt;KubePersistentVolumeFillingUp&lt;/code&gt;, on the compactor volume of our log store. The first time, we resized the volume. That bought about a month of quiet. It also made the diagnosis feel settled: more logs, not enough disk.&lt;/p&gt;
&lt;p&gt;Then the alert came back. And came back sooner.&lt;/p&gt;
&lt;p&gt;This time I opened a shell in the pod and compared two numbers:&lt;/p&gt;</description></item></channel></rss>