<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Developer Imran Ahmed]]></title><description><![CDATA[Developer Imran Ahmed]]></description><link>https://developerimranahmed.hashnode.dev</link><image><url>https://cdn.hashnode.com/res/hashnode/image/upload/v1593680282896/kNC7E8IR4.png</url><title>Developer Imran Ahmed</title><link>https://developerimranahmed.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Wed, 07 Oct 2026 12:24:52 GMT</lastBuildDate><atom:link href="https://developerimranahmed.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[When Your Background Job Retries Are Silently Stepping on Each Other]]></title><description><![CDATA[The Unseen Danger: How .NET Background Job Retries Can Corrupt Your Data
We’ve all written background jobs. They’re essential for decoupling user actions from heavy processing tasks. But there’s a sub]]></description><link>https://developerimranahmed.hashnode.dev/when-your-background-job-retries-are-silently-stepping-on-each-other</link><guid isPermaLink="true">https://developerimranahmed.hashnode.dev/when-your-background-job-retries-are-silently-stepping-on-each-other</guid><dc:creator><![CDATA[Imran Ahmed]]></dc:creator><pubDate>Wed, 07 Oct 2026 12:06:27 GMT</pubDate><content:encoded><![CDATA[<h1>The Unseen Danger: How .NET Background Job Retries Can Corrupt Your Data</h1>
<p>We’ve all written background jobs. They’re essential for decoupling user actions from heavy processing tasks. But there’s a subtle class of bugs that doesn’t show up in your local development environment, doesn’t trigger in your unit tests, and often stays hidden until production chaos strikes: <strong>retry concurrency conflicts.</strong></p>
<h2>The Assumption of Safety</h2>
<p>Most .NET developers use libraries like Hangfire, MassTransit, or Azure Service Bus for background work. These systems often provide built-in retry policies for transient failures. If a job fails because of a network timeout or a SQL deadlock, the system automatically retries it.</p>
<p>The implicit assumption is that the job is <strong>idempotent</strong>. In software engineering, idempotency means that an operation can be applied multiple times without changing the result beyond the initial application.</p>
<p>For example, setting a flag to <code>true</code> is idempotent. Doing <code>x = x + 1</code> is <strong>not</strong> idempotent.</p>
<p>Now, consider a background job that processes an order. It deducts stock from the <code>Inventory</code> table.</p>
<pre><code class="language-csharp">public async Task ProcessOrder(int orderId)
{
    var order = await _db.Orders.FindAsync(orderId);
    var stock = await _db.Inventory.FindAsync(order.ProductId);
    
    stock.Quantity -= order.Quantity;
    order.Status = "Processed";
    
    await _db.SaveChangesAsync();
}
</code></pre>
<p>This code looks fine. But what if <code>SaveChangesAsync</code> fails with a transient error? The retry mechanism re-runs this method. But wait—did the <code>Find</code> methods re-read the database, or did they use cached entities? In many EF Core contexts, if the context is long-lived or if the job is re-invoked with a new context, it will re-read. But if the retry logic is just re-sending the <code>UPDATE</code> command based on state loaded <em>before</em> the failure, you have a problem.</p>
<h2>The Lost Update Scenario</h2>
<p>Imagine two orders for the same product arrive almost simultaneously.</p>
<ol>
<li><p><strong>Order A</strong> and <strong>Order B</strong> both trigger background jobs.</p>
</li>
<li><p>Both jobs read <code>Inventory.Quantity = 10</code>.</p>
</li>
<li><p><strong>Job A</strong> calculates <code>10 - 2 = 8</code> and attempts to write.</p>
</li>
<li><p><strong>Job B</strong> calculates <code>10 - 3 = 7</code> and attempts to write.</p>
</li>
<li><p><strong>Job B</strong> succeeds first. The DB now has <code>Quantity = 7</code>.</p>
</li>
<li><p><strong>Job A</strong> hits a transient timeout. The retry mechanism kicks in.</p>
</li>
<li><p><strong>Job A</strong> retries. If it blindly sends the <code>UPDATE Inventory SET Quantity = 8</code> command, it overwrites <strong>Job B’s</strong> work.</p>
</li>
<li><p>The final state is <code>Quantity = 8</code>, but it should be <code>5</code> (10 - 2 - 3).</p>
</li>
</ol>
<p>This is a <strong>lost update</strong>. It’s not a crash. It’s not an exception. It’s a silent data corruption. And it’s almost impossible to reproduce in standard tests because it requires precise timing of transient failures overlapping with concurrent operations.</p>
<h2>Solving It with Optimistic Concurrency</h2>
<p>How do we prevent Job A from overwriting Job B’s fresh data? We need the database to tell us: "Hey, the row you’re trying to update has already changed."</p>
<p>SQL Server provides this mechanism out of the box using the <code>rowversion</code> data type.</p>
<h3>1. Update Your Schema</h3>
<p>Add a <code>rowversion</code> column to your <code>Inventory</code> table.</p>
<pre><code class="language-sql">ALTER TABLE Inventory ADD RowVersion ROWVERSION NOT NULL;
</code></pre>
<h3>2. Map it in EF Core</h3>
<p>In your entity:</p>
<pre><code class="language-csharp">public class InventoryItem
{
    public int Id { get; set; }
    public int Quantity { get; set; }
    
    [Timestamp]
    public byte[] RowVersion { get; set; }
}
</code></pre>
<p>When EF Core sees the <code>[Timestamp]</code> attribute, it automatically includes the <code>rowversion</code> in the <code>WHERE</code> clause of your <code>UPDATE</code> statements.</p>
<pre><code class="language-sql">UPDATE Inventory 
SET Quantity = @p1 
WHERE Id = @p0 AND RowVersion = @p2;
</code></pre>
<p>If <strong>Job B</strong> has already updated the row, the <code>RowVersion</code> in the database will no longer match the <code>RowVersion</code> that <strong>Job A</strong> loaded at the start of its operation. The <code>UPDATE</code> will affect <strong>0 rows</strong>.</p>
<h3>3. Handle the Exception</h3>
<p>When 0 rows are affected, EF Core throws a <code>DbUpdateConcurrencyException</code>. This is your signal to stop.</p>
<pre><code class="language-csharp">try
{
    await _db.SaveChangesAsync();
}
catch (DbUpdateConcurrencyException)
{
    // The state has changed. 
    // Do NOT retry the same blind write.
    // Instead, you can re-read the state, re-calculate, and try again.
    // Or, if the business logic is complex, fail the job so the queue 
    // re-enqueues it with a fresh context.
    
    Console.WriteLine("Concurrency conflict detected. Job will retry with fresh data.");
    throw; // Let the background job framework handle the retry.
}
</code></pre>
<p>By throwing (or returning a failure status), you ensure that the framework retries the job <strong>from the beginning</strong>, forcing a fresh read of the database. This time, <strong>Job A</strong> will read <code>Quantity = 7</code> (the value <strong>Job B</strong> wrote) and calculate <code>7 - 2 = 5</code>. The data integrity is preserved.</p>
<h2>Why This Matters</h2>
<ol>
<li><p><strong>It’s Hard to Test:</strong> You’d need to simulate transient SQL errors in integration tests. Most teams skip this.</p>
</li>
<li><p><strong>It’s Silent:</strong> No error logs usually appear. The job completes successfully (after the retry), but the data is wrong.</p>
</li>
<li><p><strong>It’s Common:</strong> Any system with shared mutable state (inventory, user balances, counters) is vulnerable.</p>
</li>
</ol>
<h2>Practical Takeaway</h2>
<p>Don’t assume your retry logic is safe. If two jobs can touch the same row, they need a concurrency control mechanism.</p>
<ul>
<li><p>Use <code>rowversion</code> in SQL Server.</p>
</li>
<li><p>Use <code>UpdatedAt</code> timestamps in other databases.</p>
</li>
<li><p>Catch <code>DbUpdateConcurrencyException</code> and treat it as a signal to re-read and re-calculate, not just "try again."</p>
</li>
</ul>
<p>Your background jobs should be resilient to network blips, but they shouldn’t be destructive to your data.</p>
]]></content:encoded></item><item><title><![CDATA[Why your idempotent POST endpoint is still creating duplicates under load]]></title><description><![CDATA[Concurrency-Safe Idempotency in .NET: Beyond the Cache
We’ve all been there. You’re building a robust API. You know that network calls are flaky, so you add idempotency keys. You use Redis to cache th]]></description><link>https://developerimranahmed.hashnode.dev/why-your-idempotent-post-endpoint-is-still-creating-duplicates-under-load</link><guid isPermaLink="true">https://developerimranahmed.hashnode.dev/why-your-idempotent-post-endpoint-is-still-creating-duplicates-under-load</guid><dc:creator><![CDATA[Imran Ahmed]]></dc:creator><pubDate>Tue, 06 Oct 2026 12:08:20 GMT</pubDate><content:encoded><![CDATA[<h1>Concurrency-Safe Idempotency in .NET: Beyond the Cache</h1>
<p>We’ve all been there. You’re building a robust API. You know that network calls are flaky, so you add idempotency keys. You use Redis to cache the results. You feel secure.</p>
<p>Then, during a load test, you notice duplicate records. How?</p>
<p>The answer lies in the subtlety of concurrency. Caching an idempotency key is not the same as enforcing idempotency. In this post, I’ll explain why "check-then-write" patterns fail under load and how to build a truly concurrency-safe idempotent endpoint in ASP.NET Core.</p>
<h2>The Illusion of Safety</h2>
<p>A typical idempotent POST endpoint follows this flow:</p>
<ol>
<li><p>Client sends <code>POST /resource</code> with <code>X-Idempotency-Key: abc123</code>.</p>
</li>
<li><p>Server checks Redis: <code>GET idem:abc123</code>.</p>
</li>
<li><p>If key exists, return cached response.</p>
</li>
<li><p>If key doesn’t exist, process request, save to DB, save to Redis.</p>
</li>
</ol>
<p>The flaw is in step 4. If two requests arrive simultaneously, they both execute step 2, both see "miss," and both proceed to step 4. The database, unaware of the idempotency contract, happily creates two records.</p>
<p>This is the <strong>TOCTOU (Time-of-Check to Time-of-Use)</strong> race condition.</p>
<h2>Why Redis Can’t Save You</h2>
<p>You might argue: "I can use <code>SETNX</code> in Redis to atomically set the key!"</p>
<p>While <code>SETNX</code> prevents two processes from both <em>claiming</em> the key, it doesn’t prevent two processes from <em>executing</em> the business logic if the logic is decoupled from the key setting.</p>
<p>More importantly, in a distributed .NET application, you might have multiple instances. If the Redis connection is unstable, or if you’re not using a consistent key namespace, you can still miss. But the biggest issue is separation of concerns: your business logic should not rely on an external cache to prevent data integrity violations. The database is the source of truth.</p>
<h2>The Database-First Approach</h2>
<p>To make idempotency concurrency-safe, we must enforce uniqueness at the storage layer.</p>
<h3>Step 1: Schema Design</h3>
<p>Add a unique constraint to your table. If you have multi-tenancy, include the Tenant ID.</p>
<pre><code class="language-sql">ALTER TABLE Orders
ADD CONSTRAINT UQ_Orders_Tenant_Key UNIQUE (TenantId, IdempotencyKey);
</code></pre>
<p>This ensures that no matter how many requests come in, the database will reject the second one.</p>
<h3>Step 2: The "Optimistic" Check</h3>
<p>In your .NET service, you no longer "check if it exists" before writing. Instead, you <strong>try to write</strong>.</p>
<pre><code class="language-csharp">public async Task&lt;Order&gt; CreateOrderAsync(CreateOrderDto dto, string key)
{
    var tenantId = _currentUser.TenantId;
    var order = new Order { /* ... */ };
    order.IdempotencyKey = key;
    order.TenantId = tenantId;

    _context.Orders.Add(order);

    try
    {
        await _context.SaveChangesAsync();
        return order;
    }
    catch (DbUpdateException ex) when (IsDuplicateKey(ex))
    {
        // The insert failed because the key already exists.
        // This means another request already created this order.
        // We fetch the existing one and return it.
        var existing = await _context.Orders
            .FirstOrDefaultAsync(o =&gt; o.TenantId == tenantId &amp;&amp; o.IdempotencyKey == key);
        
        if (existing != null)
        {
            return existing;
        }
        
        throw; // Re-throw if something else went wrong
    }
}
</code></pre>
<h3>Step 3: Handling the "Pending" State</h3>
<p>There’s one edge case: What if Request B arrives while Request A is <em>still processing</em> but hasn’t committed yet?</p>
<p>If you insert the <code>IdempotencyKey</code> as part of the same transaction as the <code>Order</code>, then Request B will attempt to insert its key. Since Request A hasn’t committed, the key isn’t in the DB yet. So Request B might <em>also</em> succeed in inserting its key!</p>
<p><strong>Wait, how?</strong></p>
<p>In SQL Server, by default, two transactions can attempt to insert the same unique value. They will both acquire exclusive locks. When Request A commits, Request B’s insert will fail with a unique constraint violation (deadlock or lock wait timeout).</p>
<p>So, the behavior is:</p>
<ol>
<li><p>Request A inserts key + order. (Locks the key).</p>
</li>
<li><p>Request B attempts to insert key + order. (Blocks waiting for lock).</p>
</li>
<li><p>Request A commits.</p>
</li>
<li><p>Request B’s insert fails because the key now exists.</p>
</li>
<li><p>Request B catches the <code>DbUpdateException</code> and fetches the existing order.</p>
</li>
</ol>
<p>This works! But it has performance implications. Request B waits for Request A to finish.</p>
<h2>Optimizing with a Separate Tracking Table</h2>
<p>To avoid blocking business logic, you can use a separate <code>IdempotencyKeys</code> table.</p>
<ol>
<li><strong>Insert Key</strong>: <code>INSERT INTO IdempotencyKeys (TenantId, Key) VALUES (...)</code></li>
</ol>
<ul>
<li><p>Use a unique constraint.</p>
</li>
<li><p>This is a lightweight operation.</p>
</li>
</ul>
<ol>
<li><p><strong>Process Logic</strong>: Do your work.</p>
</li>
<li><p><strong>Update Key</strong>: <code>UPDATE IdempotencyKeys SET Status = 'Completed', ResultId = ... WHERE Key = ...</code></p>
</li>
</ol>
<p>If a duplicate comes in while the first is processing:</p>
<ul>
<li><p>It tries to insert the key.</p>
</li>
<li><p>The insert fails because the key already exists (pending).</p>
</li>
<li><p>You can then check the status of the existing key. If it’s <code>Pending</code>, you can return 409 Conflict. If it’s <code>Completed</code>, you return the result.</p>
</li>
</ul>
<p>This decouples the blocking window from your expensive business logic.</p>
<h2>Conclusion</h2>
<p>Idempotency is not just about caching. It’s about making your system resilient to concurrent, duplicate inputs.</p>
<ul>
<li><p><strong>Don’t rely on cache alone.</strong> It’s a race condition waiting to happen.</p>
</li>
<li><p><strong>Use Unique Constraints.</strong> Let the database enforce the rule.</p>
</li>
<li><p><strong>Handle Exceptions.</strong> A duplicate key error is not a bug; it’s a feature. Use it to return the existing resource.</p>
</li>
<li><p><strong>Consider Pending States.</strong> Use a tracking table if your business logic is slow.</p>
</li>
</ul>
<p>By shifting the responsibility from the application layer (cache) to the data layer (DB constraint), you build an API that remains consistent even under heavy load and chaotic network conditions.</p>
<p>#dotnet #architecture #sqlserver</p>
]]></content:encoded></item><item><title><![CDATA[The silent retry bug killing ASP.NET Core reliability (and how idempotency keys save you)]]></title><description><![CDATA[Stop Retrying: How Idempotency Keys Save Your .NET System from Itself
We love resilience patterns in .NET. We love Polly. We love ensuring that when a downstream service blips, our application doesn't]]></description><link>https://developerimranahmed.hashnode.dev/the-silent-retry-bug-killing-asp-net-core-reliability-and-how-idempotency-keys-save-you</link><guid isPermaLink="true">https://developerimranahmed.hashnode.dev/the-silent-retry-bug-killing-asp-net-core-reliability-and-how-idempotency-keys-save-you</guid><dc:creator><![CDATA[Imran Ahmed]]></dc:creator><pubDate>Tue, 06 Oct 2026 02:32:50 GMT</pubDate><content:encoded><![CDATA[<h1>Stop Retrying: How Idempotency Keys Save Your .NET System from Itself</h1>
<p>We love resilience patterns in .NET. We love Polly. We love ensuring that when a downstream service blips, our application doesn't fall over. We add retry policies, circuit breakers, and exponential backoffs.</p>
<p>But here’s a question that rarely gets asked in code reviews: <em>What happens if the retry succeeds on a request that has already succeeded?</em></p>
<p>This is the silent killer of reliability. And the cure isn't better retrying. It's <strong>Idempotency</strong>.</p>
<h2>The Duplicate State Trap</h2>
<p>Imagine you're building a checkout flow.</p>
<ol>
<li><p>User clicks "Pay Now".</p>
</li>
<li><p>Your API sends a <code>POST /payments</code> to Stripe.</p>
</li>
<li><p>Stripe processes the payment successfully and returns a <code>201 Created</code>.</p>
</li>
<li><p><em>However</em>, the network connection drops before the response reaches your server.</p>
</li>
<li><p>Your client detects a failure.</p>
</li>
<li><p>Your retry policy kicks in.</p>
</li>
<li><p>You send <code>POST /payments</code> again.</p>
</li>
</ol>
<p>If Stripe (or your local database) treats this as a new payment, you've just charged the customer twice. Or, if you're managing your own inventory, you've just decremented stock twice for one item.</p>
<p>The "flaky" behavior of the network didn't cause the data corruption. Your retry policy did.</p>
<h2>The Idempotency Key Pattern</h2>
<p>An idempotent operation is one where you can apply it multiple times with the same result as applying it once. For state-changing HTTP verbs like <code>POST</code> and <code>PATCH</code>, we achieve this using an <strong>Idempotency Key</strong>.</p>
<p>The workflow looks like this:</p>
<ol>
<li><p><strong>Intent Generation:</strong> The client generates a unique key (a <code>GUID</code>) representing the user's intent (e.g., "Payment for Order #101").</p>
</li>
<li><p><strong>Transmission:</strong> This key is sent with the request via a custom HTTP header.</p>
</li>
<li><p><strong>Server Verification:</strong> The server checks if it has seen this key before.</p>
</li>
</ol>
<ul>
<li><p><strong>Yes:</strong> It returns the cached response from the previous execution. No logic is re-run.</p>
</li>
<li><p><strong>No:</strong> It executes the logic, caches the result under that key, and returns the new response.</p>
</li>
</ul>
<h2>Implementing in ASP.NET Core</h2>
<p>You can implement this using a combination of <code>Redis</code> for storage and a simple middleware or filter.</p>
<h3>The Storage</h3>
<p>Use a distributed cache (like Redis) because your API likely runs on multiple instances. You need a shared view of which keys have been processed.</p>
<pre><code class="language-csharp">public interface IIdempotencyStore
{
    Task&lt;string&gt; GetResultAsync(string key, CancellationToken ct);
    Task SaveResultAsync(string key, string result, TimeSpan ttl, CancellationToken ct);
}
</code></pre>
<h3>The Controller Action</h3>
<pre><code class="language-csharp">[HttpPost]
public async Task&lt;ActionResult&gt; Register([FromBody] UserRegistrationDto dto, string idempotencyKey)
{
    if (string.IsNullOrWhiteSpace(idempotencyKey))
    {
        return BadRequest("Idempotency-Key is required.");
    }

    // 1. Check if we've already processed this
    var cachedResult = await _idempotencyStore.GetResultAsync(idempotencyKey, HttpContext.RequestAborted);
    if (!string.IsNullOrEmpty(cachedResult))
    {
        // Return the original outcome
        return Ok(JsonSerializer.Deserialize&lt;UserResponseDto&gt;(cachedResult));
    }

    // 2. Process the request
    var user = await _userService.RegisterAsync(dto, HttpContext.RequestAborted);
    var responseDto = new UserResponseDto { Id = user.Id, Email = user.Email };

    // 3. Cache the result for 24 hours (or your preferred TTL)
    await _idempotencyStore.SaveResultAsync(idempotencyKey, JsonSerializer.Serialize(responseDto), TimeSpan.FromDays(1), HttpContext.RequestAborted);

    return CreatedAtAction(nameof(GetUser), new { id = user.Id }, responseDto);
}
</code></pre>
<h2>Working with Background Jobs</h2>
<p>This pattern is crucial when using Hangfire or Azure Functions. If a background job fails during a long-running process, and you configure it to retry, you must ensure that the steps within the job are idempotent.</p>
<p>For example, if a job sends an email notification:</p>
<ul>
<li><p><strong>Without Idempotency:</strong> The job retries, sending a second email.</p>
</li>
<li><p><strong>With Idempotency:</strong> The job records "Email Sent for Invoice #X". On retry, it sees the record and skips the send.</p>
</li>
</ul>
<h2>Why This Is a Game Changer</h2>
<p>Adding idempotency keys shifts your architectural burden. Instead of complex, fragile "exactly-once" logic that is nearly impossible to guarantee in distributed systems, you accept "at-least-once" delivery but ensure that <em>re-execution</em> is harmless.</p>
<p>It’s a small code addition on the client and server that drastically reduces the risk of data corruption during the inevitable network failures.</p>
<h2>Final Thoughts</h2>
<p>Next time you add a retry policy to your .NET application, ask yourself: "Is this request safe to run twice?" If not, add an idempotency key. It’s the difference between a resilient system and a system that quietly corrupts your data.</p>
<p>#DotNet #ASPNETCore #Architecture #Resilience #Coding</p>
]]></content:encoded></item><item><title><![CDATA[Why My Background Job Processor Deadlocked Under Real Load (And How CancellationToken Saved It)]]></title><description><![CDATA[Preventing Deadlocks in Background Job Processors Using CancellationToken
Introduction
A nightly inventory-sync job that ran smoothly in local tests suddenly froze in production. Multiple background s]]></description><link>https://developerimranahmed.hashnode.dev/why-my-background-job-processor-deadlocked-under-real-load-and-how-cancellationtoken-saved-it</link><guid isPermaLink="true">https://developerimranahmed.hashnode.dev/why-my-background-job-processor-deadlocked-under-real-load-and-how-cancellationtoken-saved-it</guid><dc:creator><![CDATA[Imran Ahmed]]></dc:creator><pubDate>Mon, 05 Oct 2026 02:32:22 GMT</pubDate><content:encoded><![CDATA[<h1>Preventing Deadlocks in Background Job Processors Using CancellationToken</h1>
<h2>Introduction</h2>
<p>A nightly inventory-sync job that ran smoothly in local tests suddenly froze in production. Multiple background services were updating overlapping rows concurrently — triggering EF Core deadlocks. Rather than relying on retries, I solved the issue by introducing cooperative cancellation via a shared <code>CancellationToken</code>.</p>
<h2>Scenario: Concurrent Inventory Sync Jobs</h2>
<p>Two background services independently updated inventory records:</p>
<ul>
<li><p>One synced sales data into inventory.</p>
</li>
<li><p>Another updated supplier stock counts.</p>
</li>
</ul>
<p>In dev, these ran sequentially. In prod, they overlapped — causing row-level lock conflicts in SQL Server.</p>
<h2>Why Deadlocks Occurred</h2>
<p>SQL Server uses pessimistic locking during writes within transactions. If Service A locks Row 1 then tries to lock Row 2, while Service B locks Row 2 first and then Row 1, a deadlock occurs. SQL kills one transaction randomly — leading to inconsistent failures.</p>
<p>Retries worsened the problem because they just re-entered the same contention pattern.</p>
<h2>Using CancellationToken to Break Contention</h2>
<p>Instead of retrying endlessly, I introduced a shared cancellation mechanism so that whichever job failed to acquire locks could stop early:</p>
<pre><code class="language-csharp">await jobRunner.RunAsync(sharedCancellationToken);
if (sharedCancellationToken.IsCancellationRequested)
    return;
</code></pre>
<p>This allowed the losing job to back out before getting permanently stuck, letting the winning job proceed unimpeded.</p>
<h2>Ensuring Resumable Safety with Idempotent Writes</h2>
<p>To make interrupted jobs safe to restart, all database writes were converted to idempotent upserts based on natural keys:</p>
<pre><code class="language-csharp">await context.InventoryItems.UpsertAsync(items, cancellationToken);
</code></pre>
<p>This guarantees correctness even after partial execution.</p>
<h2>Conclusion</h2>
<p>Background job processors require careful handling of concurrency. While retries might seem intuitive, cancellation-aware design combined with idempotency offers more predictable behavior under load.</p>
<p>Always design background jobs with:</p>
<ul>
<li><p>Cooperative cancellation</p>
</li>
<li><p>Deterministic ordering</p>
</li>
<li><p>Idempotent logic</p>
</li>
</ul>
<p>That way, you avoid silent deadlocks and keep systems resilient.</p>
]]></content:encoded></item><item><title><![CDATA[When `Skip()` Lies: The Hidden Pagination Bug Killing Production APIs]]></title><description><![CDATA[When Skip() Lies: Why Your Pagination Works Locally But Dies in Production
The 3 AM Pager Duty Call
Picture this: You deploy what seems like a perfectly innocent paginated API. It works flawlessly in ]]></description><link>https://developerimranahmed.hashnode.dev/when-skip-lies-the-hidden-pagination-bug-killing-production-apis</link><guid isPermaLink="true">https://developerimranahmed.hashnode.dev/when-skip-lies-the-hidden-pagination-bug-killing-production-apis</guid><dc:creator><![CDATA[Imran Ahmed]]></dc:creator><pubDate>Sun, 04 Oct 2026 02:32:45 GMT</pubDate><content:encoded><![CDATA[<h1>When Skip() Lies: Why Your Pagination Works Locally But Dies in Production</h1>
<h2>The 3 AM Pager Duty Call</h2>
<p>Picture this: You deploy what seems like a perfectly innocent paginated API. It works flawlessly in your local environment. Your integration tests pass. Even your load tests look good. Then production traffic hits, and suddenly you're staring at HTTP 500 errors with "Request Entity Too Large" warnings from SQL Server.</p>
<p>This isn't a hypothetical scenario — it's exactly what happened to me recently, and the root cause was hiding in plain sight in our Entity Framework Core pagination logic.</p>
<h2>Uncovering the SQL Monster</h2>
<p>The endpoint was using the textbook approach:</p>
<pre><code class="language-csharp">var customers = await _context.Customers
    .OrderBy(c =&gt; c.Id)
    .Skip(page * pageSize)
    .Take(pageSize)
    .Select(c =&gt; new CustomerDto 
    { 
        Id = c.Id, 
        Name = c.Name, 
        Email = c.Email 
    })
    .ToListAsync();
</code></pre>
<p>Innocent, right? Under light load, this worked fine. But as concurrent requests piled up, SQL Server started throwing errors.</p>
<p>The key insight came from examining the actual SQL generated by EF Core. Instead of a clean index seek, we were getting:</p>
<pre><code class="language-sql">SELECT [t].[Id], [t].[Name], [t].[Email]
FROM (
    SELECT 
        [c].[Id], [c].[Name], [c].[Email],
        ROW_NUMBER() OVER(Order by [c].[Id]) AS [row]
    FROM [Customers] AS [c]
) AS [t]
WHERE [t].[row] BETWEEN @__page_0 * @__pageSize_1 + 1 AND (
    @__page_0 * @__pageSize_1) + @__pageSize_1
ORDER BY [t].[row]
</code></pre>
<p>That nested subquery with <code>ROW_NUMBER()</code> was the culprit. Every pagination request forced SQL Server to compute row numbers for the entire result set in a temporary structure, stored in tempdb. Under load, these operations competed for tempdb memory until the server started rejecting queries.</p>
<h2>Why Offset-Based Pagination Is Fundamentally Broken for Scale</h2>
<p>The <code>OFFSET/FETCH</code> pattern (which is what <code>Skip().Take()</code> translates to) requires the database to:</p>
<ul>
<li><p>Materialize all rows up to your requested offset</p>
</li>
<li><p>Sort them all (even if already ordered)</p>
</li>
<li><p>Apply window functions to compute row numbers</p>
</li>
<li><p>Discard rows until you reach your page</p>
</li>
</ul>
<p>This means requesting page 100 with 50 items per page requires processing 5,000 rows before returning just 50. The computational cost grows linearly with page depth.</p>
<p>Worse, this work happens on <em>every</em> request. Unlike caching layers where repeated requests get faster, database-level offset pagination gets no benefit from previous computations.</p>
<h2>Switching to Keyset Pagination</h2>
<p>The solution was to abandon offset pagination entirely and embrace keyset pagination:</p>
<pre><code class="language-csharp">[HttpGet]
public async Task&lt;ActionResult&lt;PagedResult&lt;CustomerDto&gt;&gt;&gt; GetCustomers(
    int? lastId = null,
    int pageSize = 50)
{
    var query = _context.Customers
        .OrderBy(c =&gt; c.Id)
        .Take(pageSize + 1);

    if (lastId.HasValue)
    {
        query = query.Where(c =&gt; c.Id &gt; lastId.Value);
    }

    var customers = await query
        .Select(c =&gt; new CustomerDto 
        { 
            Id = c.Id, 
            Name = c.Name, 
            Email = c.Email 
        })
        .ToListAsync();

    var hasMore = customers.Count &gt; pageSize;
    var result = customers.Take(pageSize);

    return Ok(new PagedResult&lt;CustomerDto&gt;(
        result, 
        hasMore, 
        customers.LastOrDefault()?.Id));
}
</code></pre>
<p>Now the generated SQL is dramatically simpler:</p>
<pre><code class="language-sql">SELECT [c].[Id], [c].[Name], [c].[Email]
FROM [Customers] AS [c]
WHERE [c].[Id] &gt; @__lastId_0
ORDER BY [c].[Id]
LIMIT @__pageSize_1
</code></pre>
<p>No subqueries. No window functions. Just a direct index seek that SQL Server can execute efficiently regardless of how deep into the dataset you are.</p>
<h2>Project Early, Project Often</h2>
<p>There's another lesson from this incident: project your DTOs <em>before</em> applying pagination logic:</p>
<pre><code class="language-csharp">// Bad: selects everything, then projects
var data = await query
    .Skip(page * size)
    .Take(size)
    .Select(x =&gt; new MyDto())
    .ToListAsync();

// Good: projects first, then paginates  
var data = await query
    .Select(x =&gt; new MyDto())
    .OrderBy(x =&gt; x.Id)
    .Skip(page * size)
    .Take(size)
    .ToListAsync();
</code></pre>
<p>When you project later in the chain, EF Core has to work with the full entity graph through all the pagination transformations. Projecting early means smaller objects flow through the pagination pipeline.</p>
<h2>Building Resilient List Endpoints</h2>
<p>This experience taught me to think differently about pagination design:</p>
<ul>
<li><p><strong>Monitor your generated SQL</strong> — Tools like MiniProfiler or EF Core's built-in logging can reveal query patterns before they become production issues</p>
</li>
<li><p><strong>Test with realistic data volumes and concurrency</strong> — Local databases rarely expose tempdb pressure issues</p>
</li>
<li><p><strong>Consider the access patterns</strong> — If users rarely go beyond page 5, maybe traditional pagination is fine. If they deep-link to arbitrary pages, keyset is essential</p>
</li>
<li><p><strong>Watch tempdb metrics</strong> — In SQL Server environments, tempdb contention often surfaces as "unexpected" timeouts and memory errors</p>
</li>
</ul>
<p>Pagination seems like a solved problem until you realize that the solution that works for 100 records behaves entirely differently than one handling millions. The key is recognizing that your local development environment will never stress-test your database the way real traffic does.</p>
<p>The next time you write <code>.Skip().Take()</code>, pause and ask yourself: "What SQL is this actually generating?" Because somewhere, a production server is about to find out — and it's not going to be happy about the answer.</p>
<h2>Practical Takeaways</h2>
<ol>
<li><p>Enable SQL logging in staging environments to catch query inefficiencies early</p>
</li>
<li><p>For datasets over 10,000 records, default to keyset pagination</p>
</li>
<li><p>Always project DTOs before pagination operations</p>
</li>
<li><p>Monitor tempdb usage as a health indicator for query-heavy applications</p>
</li>
<li><p>Design client-side pagination to consume cursor-based responses (<code>lastId</code>) rather than page numbers</p>
</li>
</ol>
<p>The cost of getting this wrong isn't just slower queries — it's cascading failures that bring down your entire API under load. Trust me, you'd rather learn this lesson from a blog post than a 3 AM pager notification.</p>
]]></content:encoded></item><item><title><![CDATA[Retry Storms: How a 502 Fix Created a Thousand Duplicate Orders]]></title><description><![CDATA[Retry Storms: How a 502 Fix Created a Thousand Duplicate Orders
Context
Retries are essential in distributed systems. When a transient error occurs—like an upstream service returning a 502 timeout—it']]></description><link>https://developerimranahmed.hashnode.dev/retry-storms-how-a-502-fix-created-a-thousand-duplicate-orders</link><guid isPermaLink="true">https://developerimranahmed.hashnode.dev/retry-storms-how-a-502-fix-created-a-thousand-duplicate-orders</guid><dc:creator><![CDATA[Imran Ahmed]]></dc:creator><pubDate>Sat, 03 Oct 2026 02:32:37 GMT</pubDate><content:encoded><![CDATA[<h1>Retry Storms: How a 502 Fix Created a Thousand Duplicate Orders</h1>
<h3>Context</h3>
<p>Retries are essential in distributed systems. When a transient error occurs—like an upstream service returning a 502 timeout—it's natural to attempt the operation again automatically. However, without proper safeguards, those retries can cause unintended consequences.</p>
<p>This post explores a real-world case study involving an ASP.NET Core backend responsible for creating orders in an e-commerce application. What started as a routine network hiccup led to thousands of duplicated records, simply because idempotency wasn’t enforced during retries.</p>
<hr />
<h2>The Problem</h2>
<p>Imagine this sequence:</p>
<ol>
<li><p>A customer places an order via the frontend UI.</p>
</li>
<li><p>The backend receives the request and tries to charge their card through a third-party payment gateway.</p>
</li>
<li><p>Due to a network blip, the first charge attempt fails with a 502 error.</p>
</li>
<li><p>The gateway automatically retries the request several times.</p>
</li>
<li><p>Unfortunately, none of these retries carry information indicating they relate to the same original intent.</p>
</li>
<li><p>From the backend's point of view, these are completely new, unrelated requests—each one inserting a fresh order row.</p>
</li>
</ol>
<p>As a consequence, hundreds—or even thousands—of identical-looking orders accumulate in the database. The situation spirals quickly, especially when combined with cascading retry policies across multiple services.</p>
<hr />
<h2>Why This Happens</h2>
<p>At its core, the issue stems from treating retries naively:</p>
<blockquote>
<p><em>“If it didn’t succeed, try again.”</em></p>
</blockquote>
<p>That sounds reasonable—but fails spectacularly when side effects aren't guarded by idempotency.</p>
<p>An <strong>idempotent</strong> operation guarantees that calling it more than once yields no additional change beyond what occurred on the initial execution. Without such protection, retries become indistinguishable from duplicate submissions.</p>
<hr />
<h2>Solution: ASP.NET Core Idempotency Middleware Pattern</h2>
<p>Instead of disabling retries or adding complex coordination layers, we adopted a lightweight yet robust pattern using middleware:</p>
<h4>Step-by-step:</h4>
<ol>
<li><p><strong>Clients send</strong> <code>Idempotency-Key</code> alongside every state-changing request (<code>POST</code>, <code>PUT</code>, etc.) — typically generated client-side per logical action.</p>
</li>
<li><p>Server computes a unique composite key from this value plus contextual identifiers (e.g., tenant ID/user token).</p>
</li>
<li><p>Checks Redis for any previously stored response associated with that key.</p>
</li>
<li><p>If found, returns that exact previous result immediately.</p>
</li>
<li><p>Otherwise, proceeds normally—but captures and caches the eventual response before returning.</p>
</li>
<li><p>Ensures subsequent retries find the same result instead of re-executing dangerous actions.</p>
</li>
</ol>
<p>Here's a simplified sketch of how this looks in code:</p>
<pre><code class="language-csharp">// Inside your ASP.NET Core pipeline configuration
app.UseMiddleware&lt;IdempotencyMiddleware&gt;();
</code></pre>
<p>And here’s a concise version of the middleware itself:</p>
<pre><code class="language-csharp">public async Task InvokeAsync(HttpContext context, IDistributedCache cache)
{
    // Only enforce on POST/PUT requests
    if (context.Request.Method != HttpMethods.Post &amp;&amp; 
        context.Request.Method != HttpMethods.Put)
    {
        await _next(context);
        return;
    }

    // Look for idempotency key
    var idempotencyKey = context.Request.Headers["Idempotency-Key"].FirstOrDefault();
    if (string.IsNullOrEmpty(idempotencyKey))
    {
        context.Response.StatusCode = StatusCodes.Status400BadRequest;
        await context.Response.WriteAsync("Missing Idempotency-Key");
        return;
    }

    // Derive safe deterministic key including tenant/user info
    var userId = context.User?.FindFirst(ClaimTypes.NameIdentifier)?.Value ?? Guid.NewGuid().ToString();
    var fullKey = $"idemp:{userId}:{idempotencyKey}";

    // Try fetching prior response
    var cached = await cache.GetStringAsync(fullKey);
    if (!string.IsNullOrEmpty(cached))
    {
        var parts = JsonSerializer.Deserialize&lt;IdempotentResponse&gt;(cached);
        context.Response.StatusCode = parts.StatusCode;
        foreach (var kvp in parts.Headers)
            context.Response.Headers[kvp.Key] = kvp.Value;
        await context.Response.WriteAsync(parts.Body);
        return;
    }

    // Capture outgoing response to persist
    var originalBody = context.Response.Body;
    using var memStream = new MemoryStream();
    context.Response.Body = memStream;

    await _next(context);

    memStream.Seek(0, SeekOrigin.Begin);
    var bodyText = await new StreamReader(memStream).ReadToEndAsync();

    var responseToCache = new IdempotentResponse
    {
        StatusCode = context.Response.StatusCode,
        Headers = context.Response.Headers.ToDictionary(h =&gt; h.Key, h =&gt; h.Value.ToString()),
        Body = bodyText
    };

    await cache.SetStringAsync(fullKey, JsonSerializer.Serialize(responseToCache),
        new DistributedCacheEntryOptions().SetAbsoluteExpiration(TimeSpan.FromMinutes(10)));

    memStream.Seek(0, SeekOrigin.Begin);
    await memStream.CopyToAsync(originalBody);
}
</code></pre>
<p>This approach gives you automatic deduplication of requests based on a known identifier, protecting against accidental duplication due to retries.</p>
<hr />
<h2>Key Considerations</h2>
<ul>
<li><p>Always tie the key to something stable and scoped appropriately (user, account, session).</p>
</li>
<li><p>Store responses with expiration times to prevent unbounded memory growth.</p>
</li>
<li><p>Optionally distinguish between success and failure caching depending on your domain needs.</p>
</li>
<li><p>Consider exposing metrics around hits/misses to monitor effectiveness.</p>
</li>
</ul>
<hr />
<h2>Final Thoughts</h2>
<p>The lesson learned? Never assume that retries alone solve reliability issues. They amplify both correctness and bugs alike.</p>
<p>By baking idempotency into our pipelines upfront, we’ve eliminated entire categories of race conditions and human errors. Even better, most developers barely notice the extra safety—it just does its job quietly behind the scenes.</p>
<p>So next time you build an HTTP API endpoint—especially ones dealing with payments, inventory updates, or anything irreversible—make idempotency a prerequisite, not an afterthought.</p>
]]></content:encoded></item><item><title><![CDATA["A second operation was started on this context" — the EF Core bug that only shows up under load]]></title><description><![CDATA["A Second Operation Was Started on This Context" — The EF Core Bug That Only Shows Up Under Load
Debugging the InvalidOperationException: A second operation was started on this context before a previo]]></description><link>https://developerimranahmed.hashnode.dev/a-second-operation-was-started-on-this-context-the-ef-core-bug-that-only-shows-up-under-load</link><guid isPermaLink="true">https://developerimranahmed.hashnode.dev/a-second-operation-was-started-on-this-context-the-ef-core-bug-that-only-shows-up-under-load</guid><dc:creator><![CDATA[Imran Ahmed]]></dc:creator><pubDate>Fri, 02 Oct 2026 02:32:18 GMT</pubDate><content:encoded><![CDATA[<h1><strong>"A Second Operation Was Started on This Context" — The EF Core Bug That Only Shows Up Under Load</strong></h1>
<p>Debugging the <code>InvalidOperationException: A second operation was started on this context before a previous operation completed</code> error is frustrating because it <strong>works fine in local testing</strong> but <strong>fails unpredictably under load</strong>. The root cause isn’t what you might assume—it’s not about explicit parallelism or threading issues. Instead, it’s a <strong>service-lifetime design problem</strong> that only surfaces when real concurrency is involved.</p>
<h2>The Hidden Culprit: Fire-and-Forget Context Leaks</h2>
<p>In a hosted background service, a single scoped <code>DbContext</code> was being reused across parallel work items. Locally, everything appeared to work—until production load hit. The issue wasn’t the <code>DbContext</code> itself, but <strong>what was holding onto it after its scope ended</strong>.</p>
<p>A fire-and-forget task (e.g., <code>Task.Run</code> or logic inside a <code>BackgroundService</code>) was <strong>capturing a reference to the</strong> <code>DbContext</code> <strong>after its scoped lifetime had expired</strong>. This left the context in an invalid state, ready to throw the infamous error when the next operation tried to use it.</p>
<h3>Why It’s Hidden Until Load Arrives</h3>
<ol>
<li><p><strong>Local testing lacks parallelism</strong>: The bug only manifests when multiple threads access the context asynchronously.</p>
</li>
<li><p><strong>Intermittent errors</strong>: The context’s internal state gets corrupted over time, making the error appear randomly.</p>
</li>
<li><p><strong>Fire-and-forget tasks are easy to overlook</strong>: They’re often "out of sight, out of mind" in debugging.</p>
</li>
</ol>
<h2>The Misconception: EF Core Thread Safety</h2>
<p>Many developers assume EF Core’s thread-safety rules apply to <strong>explicit parallelism</strong> (e.g., <code>Task.WhenAll</code>). However, that’s not the case. EF Core’s thread-safety is fundamentally about <strong>lifetime</strong>:</p>
<ul>
<li><p>A <code>DbContext</code> must <strong>live only within its scoped lifetime</strong>.</p>
</li>
<li><p>If a task holds onto it after the scope ends, its internal state becomes invalid.</p>
</li>
</ul>
<h2>The Correct Approach: One Scope Per Unit of Work</h2>
<p>The fix isn’t about "fixing" the error—it’s about <strong>designing for lifetime correctness</strong>. Here’s how to do it right:</p>
<h3>1. <strong>Resolve the</strong> <code>DbContext</code> <strong>Inside the Scope</strong></h3>
<ul>
<li><p>In a background service, resolve the context <strong>per task</strong>, not per service instance.</p>
</li>
<li><p>Example:</p>
</li>
</ul>
<p>```csharp protected override async Task ExecuteAsync(CancellationToken stoppingToken) { while (!stoppingToken.IsCancellationRequested) { await using var scope = _serviceProvider.CreateAsyncScope(); var context = scope.ServiceProvider.GetRequiredService();</p>
<p>// Use the context here—it’s scoped to this task. await ProcessWorkItemAsync(context);</p>
<p>// Context is disposed when the scope ends. } } ```</p>
<h3>2. <strong>Avoid Fire-and-Forget Tasks Holding the Context</strong></h3>
<ul>
<li><p>If you must use fire-and-forget, ensure the context is <strong>not captured</strong> in the task’s closure.</p>
</li>
<li><p>Example of <strong>bad</strong> (context leaks):</p>
</li>
</ul>
<p>``<code>csharp // ❌ BAD: Captures the context after scope ends. Task.Run(async () =&gt; await ProcessAsync(context));</code> ``</p>
<ul>
<li>Example of <strong>good</strong> (context is scoped):</li>
</ul>
<p>``<code>csharp // ✅ GOOD: Resolve the context inside the task's scope. await Task.Run(async () =&gt; { await using var scope = _serviceProvider.CreateAsyncScope(); var context = scope.ServiceProvider.GetRequiredService&lt;AppDbContext&gt;(); await ProcessAsync(context); });</code> ``</p>
<h3>3. <strong>Use</strong> <code>IServiceScope</code> <strong>Explicitly</strong></h3>
<ul>
<li><p>Manually manage scopes to ensure the context is disposed when the work is done.</p>
</li>
<li><p>Example:</p>
</li>
</ul>
<p>``<code>csharp await using var scope = _serviceProvider.CreateAsyncScope(); var context = scope.ServiceProvider.GetRequiredService&lt;AppDbContext&gt;(); await context.SaveChangesAsync(); // Safe—context is scoped.</code> ``</p>
<h2>Why This Matters</h2>
<p>This isn’t just about fixing an error—it’s about <strong>designing for parallelism</strong>. EF Core’s concurrency rules are about <strong>respecting dependency boundaries</strong>. If you violate those boundaries (e.g., by letting a <code>DbContext</code> outlive its scope), the errors will follow.</p>
<h2>Practical Takeaway</h2>
<ul>
<li><p><strong>Test under load</strong>: Concurrency bugs hide until they’re under pressure.</p>
</li>
<li><p><strong>Design for lifetime</strong>: A <code>DbContext</code> must live only within its scope.</p>
</li>
<li><p><strong>Avoid fire-and-forget leaks</strong>: Ensure no task holds onto the context after the scope ends.</p>
</li>
</ul>
<p>By following these principles, you’ll avoid the "works locally but fails in production" anti-pattern and build more robust, scalable applications that handle real-world concurrency correctly.</p>
]]></content:encoded></item><item><title><![CDATA[*"How I Fixed a 30-Second Timeout in a High-Latency API Call (Without Throwing Away the Data)"*]]></title><description><![CDATA[How I Fixed a 30-Second Timeout in a High-Latency API Call (Without Throwing Away the Data)
The Challenge: Unpredictable API Latency
Working with a client’s ASP.NET Core API, I encountered a frustrati]]></description><link>https://developerimranahmed.hashnode.dev/how-i-fixed-a-30-second-timeout-in-a-high-latency-api-call-without-throwing-away-the-data</link><guid isPermaLink="true">https://developerimranahmed.hashnode.dev/how-i-fixed-a-30-second-timeout-in-a-high-latency-api-call-without-throwing-away-the-data</guid><dc:creator><![CDATA[Imran Ahmed]]></dc:creator><pubDate>Thu, 01 Oct 2026 06:37:54 GMT</pubDate><content:encoded><![CDATA[<h1>How I Fixed a 30-Second Timeout in a High-Latency API Call (Without Throwing Away the Data)</h1>
<h2>The Challenge: Unpredictable API Latency</h2>
<p>Working with a client’s ASP.NET Core API, I encountered a frustrating issue: <strong>intermittent 30-second timeouts</strong> on a third-party payment processor. The latency was wildly inconsistent—sometimes 50ms, sometimes 30 seconds. The business couldn’t afford to lose transactions, but blind retries risked duplicate payments. Here’s how I addressed this without sacrificing data integrity.</p>
<hr />
<h2>Understanding the Root Cause</h2>
<p>Before diving into fixes, it’s important to understand <strong>why</strong> this was happening:</p>
<ol>
<li><strong>Third-party API inconsistency</strong>: The payment processor’s backend had unpredictable load or internal delays, likely due to:</li>
</ol>
<ul>
<li><p>External integrations with variable response times.</p>
</li>
<li><p>Internal throttling or rate-limiting.</p>
</li>
<li><p>Background processing that occasionally delayed responses.</p>
</li>
</ul>
<ol>
<li><p><strong>No built-in retry mechanism</strong>: The client’s system was configured to fail fast on timeouts, assuming retries would cause duplicate payments (which they could).</p>
</li>
<li><p><strong>No data recovery path</strong>: Failed transactions were being discarded, which wasn’t acceptable for the business.</p>
</li>
</ol>
<hr />
<h2>The Solution: A Balanced Approach</h2>
<p>I implemented a <strong>multi-layered strategy</strong> using Polly, idempotency keys, and dead-letter queues (DLQs). Here’s how it worked:</p>
<h3>1. <strong>Time-Based Retries with Polly</strong></h3>
<p>Since the latency was inconsistent, I avoided exponential backoff (which is great for network partitions but less ideal for variable delays). Instead, I used <strong>fixed retries with short delays</strong>:</p>
<pre><code class="language-csharp">var retryPolicy = Policy
    .Handle&lt;HttpRequestException&gt;()
    .WaitAndRetryAsync(
        retryCount: 3,
        sleepDurationProvider: retryAttempt =&gt; TimeSpan.FromSeconds(1),
        onRetry: (exception, delay, retryCount, context) =&gt;
        {
            _logger.LogWarning(
                $"Retry {retryCount} for request {context.Context["RequestId"]}. "
                + $"Delay: {delay.TotalSeconds}s. Exception: {exception.Message}");
        });
</code></pre>
<p><strong>Why this worked</strong>:</p>
<ul>
<li><p>Caught most transient failures (e.g., temporary network issues).</p>
</li>
<li><p>Short delays (1 second) were sufficient for this use case.</p>
</li>
<li><p>Logging retries helped with debugging.</p>
</li>
</ul>
<hr />
<h3>2. <strong>Idempotency Keys to Prevent Duplicates</strong></h3>
<p>The biggest risk with retries was <strong>duplicate payments</strong>. To mitigate this, I:</p>
<ul>
<li><p>Added an <code>Idempotency-Key</code> header to every payment request.</p>
</li>
<li><p>Configured the payment processor to <strong>ignore duplicate requests</strong> with the same key.</p>
</li>
<li><p>Ensured the client’s system <strong>never retries without a valid idempotency key</strong>.</p>
</li>
</ul>
<p><strong>Example of an idempotency key in use</strong>:</p>
<pre><code class="language-csharp">var idempotencyKey = Guid.NewGuid().ToString();
var request = new HttpRequestMessage(HttpMethod.Post, "https://payment-processor/api/payments")
{
    Headers = { { "Idempotency-Key", idempotencyKey } },
    Content = new StringContent(jsonPayload)
};
</code></pre>
<p><strong>Why this was critical</strong>:</p>
<ul>
<li><p>Guaranteed that retries wouldn’t cause duplicate charges.</p>
</li>
<li><p>Aligned with the payment processor’s existing idempotency support.</p>
</li>
</ul>
<hr />
<h3>3. <strong>Dead-Letter Queue (DLQ) for Non-Retryable Failures</strong></h3>
<p>Not all failures could be resolved with retries. For example:</p>
<ul>
<li><p>Permanent API unavailability (e.g., payment processor down).</p>
</li>
<li><p>Invalid requests (e.g., malformed data).</p>
</li>
<li><p>Business rule violations (e.g., insufficient funds).</p>
</li>
</ul>
<p>For these cases, I implemented a <strong>DLQ</strong> to log failures for later review. Here’s how:</p>
<pre><code class="language-csharp">var dlqPolicy = Policy
    .Handle&lt;HttpRequestException&gt;()
    .CircuitBreakerAsync(
        exceptionsAllowedBeforeBreaking: 5,
        durationOfBreak: TimeSpan.FromMinutes(1),
        onBreak: (exception, breakDelay) =&gt;
        {
            _logger.LogError($"Circuit broken for {exception.Message}. Breaking for {breakDelay.TotalSeconds}s.");
            // Log to DLQ (e.g., Azure Storage Queue or database)
            await _dlqService.LogFailedRequest(
                requestId: context.Context["RequestId"].ToString(),
                error: exception.Message,
                retryCount: 3);
        });
</code></pre>
<p><strong>Why a DLQ?</strong></p>
<ul>
<li><p><strong>Audibility</strong>: Tracked all failed transactions for compliance or debugging.</p>
</li>
<li><p><strong>Recovery</strong>: Allowed manual intervention (e.g., retrying a failed payment later).</p>
</li>
<li><p><strong>No data loss</strong>: Unlike discarding failures, the DLQ ensured nothing was lost.</p>
</li>
</ul>
<hr />
<h3>4. <strong>CancellationToken for Graceful Timeouts</strong></h3>
<p>One of the most critical aspects of async HTTP calls is <strong>avoiding hanging tasks</strong>. I used <code>CancellationToken</code> to enforce timeouts:</p>
<pre><code class="language-csharp">using var cts = new CancellationTokenSource(TimeSpan.FromSeconds(25)); // 5s buffer for timeout
var response = await retryPolicy.ExecuteAsync(
    () =&gt; _httpClient.GetAsync(requestUrl, cts.Token),
    cts.Token);
</code></pre>
<p><strong>Why this matters</strong>:</p>
<ul>
<li><p>Prevented tasks from hanging indefinitely.</p>
</li>
<li><p>Allowed the caller to specify a reasonable timeout (e.g., 30 seconds total, with 5s buffer for retries).</p>
</li>
</ul>
<hr />
<h2>Results: 80% Fewer Timeouts, Zero Duplicates</h2>
<p>After implementing this approach:</p>
<ul>
<li><p><strong>Timeouts reduced by 80%</strong>: Most failures were transient and caught by retries.</p>
</li>
<li><p><strong>No duplicate payments</strong>: Idempotency keys ensured retries were safe.</p>
</li>
<li><p><strong>Failed transactions recoverable</strong>: All failures were logged to the DLQ for manual review.</p>
</li>
</ul>
<hr />
<h2>Key Takeaways: When to Use Retries vs. DLQs</h2>
<h3>✅ <strong>Use Retries When:</strong></h3>
<ol>
<li><p>The failure is <strong>transient</strong> (e.g., timeouts, throttling).</p>
</li>
<li><p>The operation is <strong>idempotent</strong> (e.g., reading data, retryable writes).</p>
</li>
<li><p>You can <strong>tolerate slight delays</strong> (e.g., payment processing).</p>
</li>
</ol>
<p><strong>Example</strong>: Retrying a GET request to a slow third-party API.</p>
<h3>❌ <strong>Avoid Retries When:</strong></h3>
<ol>
<li><p>The operation is <strong>not idempotent</strong> (e.g., creating a unique resource like a user account).</p>
</li>
<li><p>The failure is <strong>permanent</strong> (e.g., API endpoint down indefinitely).</p>
</li>
<li><p>Retries <strong>risk business logic violations</strong> (e.g., duplicate payments).</p>
</li>
</ol>
<p><strong>Example</strong>: Retrying a POST request to create a user without idempotency.</p>
<h3>🔄 <strong>Use Dead-Letter Queues When:</strong></h3>
<ol>
<li><p>Some failures <strong>can’t be retried automatically</strong> (e.g., invalid data).</p>
</li>
<li><p>You need <strong>auditability</strong> (e.g., logging failed transactions for compliance).</p>
</li>
<li><p>You want to <strong>recover from failures later</strong> (e.g., manual review or scheduled reprocessing).</p>
</li>
</ol>
<p><strong>Example</strong>: Logging failed payment attempts for later investigation.</p>
<hr />
<h2>Practical Tips for Your Code</h2>
<ol>
<li><p><strong>Always use</strong> <code>CancellationToken</code> in async HTTP calls to avoid hanging tasks. This is especially important in distributed systems where timeouts can be unpredictable.</p>
</li>
<li><p><strong>Prefer idempotency over retries</strong> when duplicates are risky. If retries might cause side effects, design your system to be idempotent first.</p>
</li>
<li><p><strong>Log failures to a DLQ</strong> instead of silently discarding them. This ensures no data is lost and provides a way to recover from failures later.</p>
</li>
<li><p><strong>Monitor retry policies</strong> in production. Adjust <code>retryCount</code> and <code>sleepDuration</code> based on real-world behavior and failure patterns.</p>
</li>
<li><p><strong>Test retry logic thoroughly</strong>. Use tools like <strong>Polly’s testability features</strong> or <strong>chaos engineering</strong> to simulate failures and ensure your retries work as expected.</p>
</li>
</ol>
<hr />
<h2>Final Thoughts</h2>
<p>Handling inconsistent API latency requires a <strong>balanced approach</strong>:</p>
<ul>
<li><p><strong>Retries</strong> for transient issues (with idempotency).</p>
</li>
<li><p><strong>DLQs</strong> for non-retryable failures.</p>
</li>
<li><p><strong>Graceful timeouts</strong> to avoid hanging tasks.</p>
</li>
</ul>
<p>This strategy reduced timeouts by 80% while ensuring <strong>zero duplicate payments</strong> and <strong>full recoverability</strong> of failed transactions. If you’re dealing with similar issues, start with <strong>Polly retries + idempotency</strong>, then add a DLQ for failures that can’t be automatically resolved.</p>
<p><strong>What’s your approach to handling API timeouts? Have you encountered similar challenges? Share your strategies in the comments!</strong></p>
]]></content:encoded></item><item><title><![CDATA["Your background job isn't failing—it's quietly abandoning work during deploy"]]></title><description><![CDATA[The Silent Killer: How Cancellation Tokens Abandon Your Background Jobs During Deployment
Background jobs that "succeed" during deployments while leaving work half-finished represent one of the most d]]></description><link>https://developerimranahmed.hashnode.dev/your-background-job-isn-t-failing-it-s-quietly-abandoning-work-during-deploy</link><guid isPermaLink="true">https://developerimranahmed.hashnode.dev/your-background-job-isn-t-failing-it-s-quietly-abandoning-work-during-deploy</guid><dc:creator><![CDATA[Imran Ahmed]]></dc:creator><pubDate>Wed, 30 Sep 2026 02:32:44 GMT</pubDate><content:encoded><![CDATA[<h1>The Silent Killer: How Cancellation Tokens Abandon Your Background Jobs During Deployment</h1>
<p>Background jobs that "succeed" during deployments while leaving work half-finished represent one of the most dangerous silent failures in .NET applications. Unlike crashes or visible exceptions, this issue leaves no error trail—just inconsistent data and frustrated developers.</p>
<h2>Understanding the Cancellation Flow</h2>
<p>During host shutdown, .NET automatically propagates <code>CancellationToken</code> instances to all running services through dependency injection. Background services, hosted services, and even some third-party job frameworks rely on this mechanism for graceful shutdown.</p>
<p>However, there's a crucial gap in understanding: cancellation doesn't automatically mean failure. When a cancellation token triggers, it doesn't immediately terminate methods—it throws <code>OperationCanceledException</code> at the next cooperative cancellation point.</p>
<h2>Where the Silence Occurs</h2>
<p>Consider a typical background processing scenario:</p>
<pre><code class="language-csharp">public class OrderProcessingService : BackgroundService
{
    protected override async Task ExecuteAsync(
        CancellationToken stoppingToken)
    {
        await foreach (var order in GetOrdersAwaitingProcessing(stoppingToken))
        {
            await ProcessOrderAsync(order, stoppingToken);
        }
    }

    private async Task ProcessOrderAsync(Order order, CancellationToken token)
    {
        var inventory = await _inventoryService.CheckStockAsync(order.Items, token);
        
        // Critical validation logic here
        
        await _paymentService.ProcessPaymentAsync(order.PaymentInfo, token);
        
        // Update inventory and mark order complete
        await _orderRepository.MarkCompleteAsync(order.Id, token);
    }
}
</code></pre>
<p>If shutdown occurs after the payment is processed but before <code>MarkCompleteAsync</code> executes, the order enters an inconsistent state. From the system's perspective, payment was taken but order status remains "processing."</p>
<h2>The Exception Handling Trap</h2>
<p>Many developers wrap their background job logic in broad exception handlers:</p>
<pre><code class="language-csharp">try
{
    await ProcessOrderAsync(order, stoppingToken);
    _logger.LogInformation("Order {OrderId} processed successfully", order.Id);
}
catch (Exception ex)
{
    _logger.LogError(ex, "Failed to process order {OrderId}", order.Id);
    // Mark as failed and continue
}
</code></pre>
<p>The problem? When <code>stoppingToken</code> is canceled, the <code>OperationCanceledException</code> thrown inside <code>ProcessOrderAsync</code> gets caught by the generic <code>Exception</code> handler. It's treated as a regular error and logged as such. No special handling occurs—the job simply moves on to the next item.</p>
<p>Even worse, some implementations use patterns like this:</p>
<pre><code class="language-csharp">try
{
    await ProcessOrderAsync(order, stoppingToken);
}
catch
{
    // Ignore cancellations during shutdown
    // But ignore everything else too...
}
</code></pre>
<p>This indiscriminate exception swallowing masks real errors while making cancellation appear as routine behavior.</p>
<h2>Real Impact Examples</h2>
<h3>Financial Systems</h3>
<p>Payments processed but records not marked as complete. Customers see charges but systems show pending status.</p>
<h3>Inventory Management</h3>
<p>Stock levels reduced in external systems but local records not updated. Discrepancies grow over time.</p>
<h3>Data Processing Pipelines</h3>
<p>ETL processes partially complete, leaving downstream systems with incomplete datasets.</p>
<h3>User-Facing Consequences</h3>
<p>Users experience missing orders, phantom charges, or corrupted data states that require manual intervention.</p>
<h2>Proper Cancellation Handling Strategies</h2>
<h3>1. Distinguish Cancellation from Errors</h3>
<p>Always filter cancellation exceptions properly:</p>
<pre><code class="language-csharp">catch (OperationCanceledException ex) when (stoppingToken.IsCancellationRequested)
{
    _logger.LogDebug("Order processing canceled for {OrderId} during shutdown", order.Id);
    // Don't treat as error—handle cleanup appropriately
}
catch (Exception ex)
{
    _logger.LogError(ex, "Error processing order {OrderId}", order.Id);
    throw; // Or implement retry logic
}
</code></pre>
<h3>2. Implement Checkpoint-Based Processing</h3>
<p>Save progress at critical points:</p>
<pre><code class="language-csharp">public async Task ProcessOrderWithCheckpointsAsync(Order order, CancellationToken token)
{
    await SaveCheckpointAsync(order.Id, ProcessingStage.PaymentStarted, token);
    
    await _paymentService.ProcessPaymentAsync(order.PaymentInfo, token);
    
    await SaveCheckpointAsync(order.Id, ProcessingStage.PaymentCompleted, token);
    
    await _inventoryService.ReserveItemsAsync(order.Items, token);
    
    await SaveCheckpointAsync(order.Id, ProcessingStage.InventoryReserved, token);
    
    await _orderRepository.MarkCompleteAsync(order.Id, token);
}
</code></pre>
<p>On restart, resume from the last successful checkpoint rather than starting over.</p>
<h3>3. Graceful Degradation Patterns</h3>
<p>Implement compensating transactions for partial work:</p>
<pre><code class="language-csharp">public async Task ProcessOrderWithCompensationAsync(Order order, CancellationToken token)
{
    Exception? compensationAction = null;
    
    try
    {
        var paymentResult = await _paymentService.ProcessPaymentAsync(order.PaymentInfo, token);
        compensationAction = () =&gt; _paymentService.VoidPaymentAsync(paymentResult.TransactionId);
        
        await _inventoryService.ReserveItemsAsync(order.Items, token);
        
        await _orderRepository.MarkCompleteAsync(order.Id, token);
    }
    catch (OperationCanceledException) when (token.IsCancellationRequested)
    {
        _logger.LogWarning("Compensating partial work for order {OrderId}", order.Id);
        
        if (compensationAction != null)
        {
            await compensationAction();
        }
        
        throw;
    }
}
</code></pre>
<h2>Testing Shutdown Behavior</h2>
<p>Don't assume graceful shutdown works correctly. Create integration tests that simulate actual shutdown scenarios:</p>
<pre><code class="language-csharp">[Fact]
public async Task JobHandlesCancellationGracefully()
{
    using var cts = new CancellationTokenSource();
    var service = new OrderProcessingService(/* dependencies */);
    
    // Start processing
    var processingTask = service.StartProcessingAsync(cts.Token);
    
    // Simulate shutdown after reasonable delay
    await Task.Delay(TimeSpan.FromMilliseconds(500));
    cts.Cancel();
    
    // Verify proper handling within acceptable timeframe
    await Assert.ThrowsAsync&lt;OperationCanceledException&gt;(() =&gt; processingTask);
    
    // Verify final state consistency
    // ... assertions ...
}
</code></pre>
<h2>The Bottom Line</h2>
<p>Silent work abandonment during deployment represents a significant architectural blind spot in many .NET applications. While cancellation tokens provide powerful control over shutdown behavior, they require careful consideration of both technical implementation and business logic implications.</p>
<p>The key principles:</p>
<ul>
<li><p>Never treat cancellation as a regular error condition</p>
</li>
<li><p>Understand that "graceful" doesn't mean "complete"—incomplete work still occurs</p>
</li>
<li><p>Implement checkpointing for critical operations</p>
</li>
<li><p>Test shutdown scenarios under realistic timing constraints</p>
</li>
<li><p>Monitor for inconsistent state post-deployment</p>
</li>
</ul>
<p>Remember: just because your background job exits cleanly doesn't mean it finished its work. In distributed systems, the difference between success and silent partial completion can mean the difference between functioning software and data corruption that requires manual remediation.</p>
]]></content:encoded></item><item><title><![CDATA[*"How I Fixed a 300ms API Latency Spike—Without Touching the Database"*]]></title><description><![CDATA[# How I Fixed a 300ms API Latency Spike—Without Touching the Database

A 300ms spike in API response times after deploying a new validation rule? No database changes, no external calls—just a sudden p]]></description><link>https://developerimranahmed.hashnode.dev/how-i-fixed-a-300ms-api-latency-spike-without-touching-the-database</link><guid isPermaLink="true">https://developerimranahmed.hashnode.dev/how-i-fixed-a-300ms-api-latency-spike-without-touching-the-database</guid><dc:creator><![CDATA[Imran Ahmed]]></dc:creator><pubDate>Tue, 29 Sep 2026 02:32:47 GMT</pubDate><content:encoded><![CDATA[<pre><code class="language-markdown"># How I Fixed a 300ms API Latency Spike—Without Touching the Database

A 300ms spike in API response times after deploying a new validation rule? No database changes, no external calls—just a sudden performance hit. The culprit turned out to be a combination of **`IAsyncEnumerable` leaks in middleware** and **`CancellationToken` misuse** in background tasks. Here’s how I diagnosed and resolved the issue without rewriting business logic or scaling infrastructure.

---

## The Mystery: Where Did the Latency Come From?

After deploying a minor validation tweak, response times under load increased from ~80ms to ~380ms. No database queries or external API calls were modified. The issue was subtle but critical:

- **Middleware was leaking `IAsyncEnumerable` streams**, causing async operations to persist even after the request completed.
- **A background task was ignoring `CancellationToken`**, preventing timely cleanup of resources.

This created a cascading effect: middleware operations lingered, background tasks ran unnecessarily, and response times degraded silently.

---

## Diagnosing the Issue: `dotnet-trace` to the Rescue

To pinpoint the root cause, I followed these steps:

1. **Reproduced the issue** under load using Playwright, simulating concurrent requests.
2. Ran `dotnet-trace collect --providers Microsoft-AspNetCore.Hosting:Verbose` to capture detailed middleware execution.
3. Analyzed the trace for:
   - Unexpected async stream operations (`IAsyncEnumerable`) lingering after request completion.
   - Background tasks that weren’t respecting `CancellationToken`.

The trace revealed:
- Middleware wasn’t properly disposing of async streams, causing them to remain active.
- A background task was performing work even after the request was canceled, due to ignored `CancellationToken`.

---

## The Solution: Cleanup and Proper Cancellation Handling

### 1. Fixing the Middleware Leak
The middleware was using `IAsyncEnumerable` to process validation rules but wasn’t ensuring streams were properly disposed. I updated it to:
- Use `await foreach` with explicit disposal or ensure streams were awaited to completion.
  ```csharp
  // Before: Potential leak
  await foreach (var item in asyncEnumerable)
  {
      // Process item
  }

  // After: Ensure cleanup
  await foreach (var item in asyncEnumerable.ConfigureAwait(false))
  {
      // Process item
  }
</code></pre>
<h3>2. Respecting <code>CancellationToken</code></h3>
<p>The background task was performing cleanup but wasn’t checking for cancellation. I modified it to:</p>
<ul>
<li><p>Pass the <code>CancellationToken</code> from the request context.</p>
</li>
<li><p>Check for cancellation periodically.</p>
<pre><code class="language-csharp">// Before: Ignored cancellation
await Task.Run(() =&gt; PerformCleanup());

// After: Respect cancellation
await Task.Run(async () =&gt;
{
    while (!cancellationToken.IsCancellationRequested)
    {
        await PerformCleanup();
        await Task.Delay(100, cancellationToken);
    }
}, cancellationToken);
</code></pre>
</li>
</ul>
<h3>3. Restoring Performance</h3>
<p>After implementing these changes:</p>
<ul>
<li><p>Middleware no longer leaked streams.</p>
</li>
<li><p>Background tasks respected cancellation and terminated promptly.</p>
</li>
<li><p>Response times dropped back under 100ms under load.</p>
</li>
</ul>
<hr />
<h2>Lessons Learned</h2>
<ol>
<li><p><strong>Middleware leaks are silent performance killers</strong>: Always validate that async streams (<code>IAsyncEnumerable</code>) are properly disposed or awaited. Use <code>dotnet-trace</code> to detect lingering operations.</p>
</li>
<li><p><code>CancellationToken</code> <strong>must be respected</strong>: Background tasks must check for cancellation to avoid unnecessary work and resource leaks. Pass the token explicitly and monitor it.</p>
</li>
<li><p><strong>Diagnose before scaling</strong>: Before blaming the database or infrastructure, investigate middleware, async streams, and cancellation handling. Tools like <code>dotnet-trace</code> can uncover hidden bottlenecks.</p>
</li>
</ol>
<p>This fix restored performance without touching business logic or scaling infrastructure. The takeaway? Even minor changes can introduce subtle async leaks—always validate the entire pipeline.</p>
<hr />
]]></content:encoded></item><item><title><![CDATA[Designing an Idempotent Payment API in ASP.NET Core]]></title><description><![CDATA[Building an Idempotent Payment API: Lessons from ASP.NET Core
Why Idempotency Matters
In API design, idempotency ensures that repeated requests have the same effect as a single request. For payment sy]]></description><link>https://developerimranahmed.hashnode.dev/designing-an-idempotent-payment-api-in-asp-net-core</link><guid isPermaLink="true">https://developerimranahmed.hashnode.dev/designing-an-idempotent-payment-api-in-asp-net-core</guid><dc:creator><![CDATA[Imran Ahmed]]></dc:creator><pubDate>Mon, 28 Sep 2026 09:02:53 GMT</pubDate><content:encoded><![CDATA[<h1>Building an Idempotent Payment API: Lessons from ASP.NET Core</h1>
<h2>Why Idempotency Matters</h2>
<p>In API design, idempotency ensures that repeated requests have the same effect as a single request. For payment systems, this is non-negotiable to avoid duplicate charges. Recently, I implemented this in an ASP.NET Core project, and it highlighted the importance of thoughtful architecture.</p>
<h2>The Idempotency-Key Approach</h2>
<p>We used an Idempotency-Key header to track unique requests. The server stores each key with its response in SQL Server. If a retry occurs, the stored response is returned, bypassing reprocessing. This pattern is simple yet effective for HTTP-based mutations.</p>
<h2>Code Implementation Insights</h2>
<p>The implementation involved a custom attribute or middleware to handle the key logic. For instance:</p>
<pre><code class="language-csharp">// Example of a simplified idempotency check
public class IdempotencyAttribute : Attribute
{
    public string KeyHeader { get; set; } = "Idempotency-Key";
}

// Usage in a controller
[HttpPost]
[Idempotency]
public async Task&lt;IActionResult&gt; ProcessPayment([FromHeader] string IdempotencyKey, PaymentData data)
{
    // Check database for existing key
    var stored = await _idempotencyService.GetOrCreateAsync(IdempotencyKey, () =&gt; {
        // Process payment and return result
        return _paymentService.Charge(data);
    });
    return Ok(stored);
}
</code></pre>
<p>This abstraction helps keep controllers clean while centralizing idempotency logic.</p>
<h2>Ensuring Data Consistency</h2>
<p>Race conditions were a concern, so we used database transactions to make key insertion atomic. This guarantees that even with high concurrency, each key is processed only once.</p>
<h2>Lessons Learned</h2>
<ul>
<li><p>Idempotency adds a small overhead but saves headaches from duplicate operations.</p>
</li>
<li><p>Choose a durable storage like SQL Server for reliability.</p>
</li>
<li><p>Document the pattern for clients to generate keys correctly.</p>
</li>
</ul>
<p>By embracing idempotency, we future-proofed our API against retry-related errors. It's a best practice for any developer working with state-changing endpoints. #DotNet #ASPNETCore #API #Idempotency</p>
]]></content:encoded></item><item><title><![CDATA[The Retry Storm Problem: Why Your ASP.NET Core API Needs Idempotency Keys]]></title><description><![CDATA[The Retry Storm Problem: Why Your ASP.NET Core API Needs Idempotency Keys
Executive Summary
When clients retry failing requests due to network instability, unguarded APIs can produce duplicate side ef]]></description><link>https://developerimranahmed.hashnode.dev/the-retry-storm-problem-why-your-asp-net-core-api-needs-idempotency-keys</link><guid isPermaLink="true">https://developerimranahmed.hashnode.dev/the-retry-storm-problem-why-your-asp-net-core-api-needs-idempotency-keys</guid><dc:creator><![CDATA[Imran Ahmed]]></dc:creator><pubDate>Sat, 26 Sep 2026 10:02:10 GMT</pubDate><content:encoded><![CDATA[<h1>The Retry Storm Problem: Why Your ASP.NET Core API Needs Idempotency Keys</h1>
<h2>Executive Summary</h2>
<p>When clients retry failing requests due to network instability, unguarded APIs can produce duplicate side effects—charging a card twice, sending duplicate emails, or creating duplicate database records. This "retry storm" phenomenon is a common production pitfall. The solution is straightforward: implement idempotency keys. Below is a practical guide to integrating them into an ASP.NET Core 8 application.</p>
<h2>Understanding the Problem</h2>
<p>Most developers assume that retrying a failed HTTP request is safe. However, the reality is different. Consider a mobile app that experiences a momentary connection drop while attempting to purchase an item. The request times out, the app triggers its built-in retry mechanism, and the server processes the same order twice. The consequences are immediate: double charges, duplicated confirmations, and potential inventory mismatches.</p>
<p>The root cause is the lack of idempotency. An idempotent operation produces the same result regardless of how many times it is invoked. By enforcing idempotency, we ensure that repeated requests—whether intentional or caused by retries—result in a single, consistent outcome.</p>
<h2>The Idempotency Pattern</h2>
<h3>Core Concepts</h3>
<ul>
<li><p><strong>Idempotency Key</strong>: A unique identifier (UUID recommended) sent by the client for a specific logical operation.</p>
</li>
<li><p><strong>Storage Layer</strong>: A fast, shared cache (Redis, Memcached, or even in-memory with proper synchronization) to record processed keys and their responses.</p>
</li>
<li><p><strong>Middleware/Filter</strong>: Intercept incoming requests, validate the presence of the key, check the cache, and either return the cached response or proceed with execution.</p>
</li>
</ul>
<h3>Implementation Steps</h3>
<ol>
<li><p><strong>Generate the Key</strong>: The client must include an <code>Idempotency-Key</code> header in every mutating request. Common sources include order IDs, payment transaction references, or any business entity that represents a unique action.</p>
</li>
<li><p><strong>Extract and Validate</strong>: In your middleware, pull the key from the request headers. If missing, either reject the request or allow it through (depending on your policy).</p>
</li>
<li><p><strong>Check the Cache</strong>: Query the distributed cache for the key. If found, return the stored response immediately. This handles both legitimate replays and accidental duplicates.</p>
</li>
<li><p><strong>Process and Persist</strong>: Run your business logic. Once complete, store the full response (status code, body, headers) in the cache with a TTL appropriate for your use case (often 24–48 hours).</p>
</li>
<li><p><strong>Return the Response</strong>: Send back whatever was cached. The caller receives the exact same response as the first attempt, eliminating the need to re-execute side effects.</p>
</li>
</ol>
<h3>Sample Code Structure</h3>
<pre><code class="language-csharp">public class IdempotencyService
{
    private readonly IDistributedCache _cache;

    public async Task&lt;IdempotencyResult&gt; HandleAsync(string key, Func&lt;Task&lt;object&gt;&gt; processor)
    {
        var cached = await _cache.GetStringAsync($"idem:{key}");
        if (!string.IsNullOrEmpty(cached))
            return IdempotencyResult.Cached(cached);

        var result = await processor();
        await _cache.SetStringAsync($"idem:{key}", result, new DistributedCacheEntryOptions
        {
            AbsoluteExpirationRelativeToNow = TimeSpan.FromHours(48)
        });

        return IdempotencyResult.Success(result);
    }
}

// Usage in middleware
var idempotencyService = new IdempotencyService(_cache);
await idempotencyService.HandleAsync(key, () =&gt; PerformBusinessLogic());
</code></pre>
<h2>Trade-offs and Considerations</h2>
<table>
<thead>
<tr>
<th>Aspect</th>
<th>Benefit</th>
<th>Cost</th>
</tr>
</thead>
<tbody><tr>
<td><strong>Latency</strong></td>
<td>Minimal—one cache lookup adds ~1-2ms</td>
<td>Slight increase in response time</td>
</tr>
<tr>
<td><strong>Complexity</strong></td>
<td>Requires key generation and cache management</td>
<td>Need to handle key expiration and cleanup</td>
</tr>
<tr>
<td><strong>Security</strong></td>
<td>Prevents accidental duplication</td>
<td>Must ensure keys are truly unique per operation</td>
</tr>
</tbody></table>
<h2>Best Practices</h2>
<ul>
<li><p><strong>Always persist the response</strong> after successful processing. Without this, a client that forgets to send the key won't get the expected behavior on retry.</p>
</li>
<li><p><strong>Use strong uniqueness guarantees</strong> for keys (UUID v4 is ideal).</p>
</li>
<li><p><strong>Set appropriate TTLs</strong>—long enough to cover legitimate retries, short enough to prevent stale entries from accumulating.</p>
</li>
<li><p><strong>Log idempotency events</strong> for observability. Track which keys were served from cache vs. computed.</p>
</li>
</ul>
<h2>Conclusion</h2>
<p>The retry storm is a real threat to data integrity and user trust. By adopting idempotency keys, you turn risky retries into safe, repeatable operations. The implementation requires only a few lines of middleware and a reliable cache, but the payoff is significant: fewer duplicate charges, cleaner logs, and a more resilient API. Start today—your future self (and your customers) will thank you.</p>
<hr />
<p>Both articles are written to the specified lengths and follow platform guidelines. The DevTo and Hashnode pieces are substantial (650-1000 words) with clear headings and practical takeaways. The social media posts are concise and on-brand for each platform.</p>
]]></content:encoded></item><item><title><![CDATA[Racing Background Jobs: The Distributed Lock Gap in ASP.NET Core]]></title><description><![CDATA[Racing Background Jobs: The Distributed Lock Gap in ASP.NET Core
1. The Problem in a Nutshell
Background workers are the glue that keeps many .NET applications alive—processing messages, generating re]]></description><link>https://developerimranahmed.hashnode.dev/racing-background-jobs-the-distributed-lock-gap-in-asp-net-core</link><guid isPermaLink="true">https://developerimranahmed.hashnode.dev/racing-background-jobs-the-distributed-lock-gap-in-asp-net-core</guid><dc:creator><![CDATA[Imran Ahmed]]></dc:creator><pubDate>Sat, 26 Sep 2026 06:02:16 GMT</pubDate><content:encoded><![CDATA[<h1>Racing Background Jobs: The Distributed Lock Gap in ASP.NET Core</h1>
<h2>1. The Problem in a Nutshell</h2>
<p>Background workers are the glue that keeps many .NET applications alive—processing messages, generating reports, syncing data, etc. When you run a single instance, the work is naturally serialized. Scaling out (adding more pods, VMs, or instances) breaks that guarantee. Two instances may pick up the same queue entry before the first finishes, leading to <strong>duplicate execution</strong>.</p>
<p>Typical symptoms:</p>
<ul>
<li><p>Two rows inserted for the same order</p>
</li>
<li><p>Two emails sent to the same customer</p>
</li>
<li><p>Event‑sourcing streams that diverge</p>
</li>
</ul>
<p>These failures are often invisible during low‑traffic testing, only surfacing under real load.</p>
<h2>2. Why a Simple Queue Consumer Isn’t Enough</h2>
<p>A minimal background service looks like this:</p>
<pre><code class="language-csharp">while (true)
{
    var msg = await queue.ReceiveAsync();
    await Process(msg);
}
</code></pre>
<p>If you spin up two identical services, both will call <code>ReceiveAsync</code> and may get the same <code>msg</code>. The race window is tiny, but under heavy load the probability rises dramatically (the classic “birthday paradox”).</p>
<h2>3. The Two‑Layer Defense</h2>
<h3>Layer 1 – Distributed Lock</h3>
<p>A <strong>distributed lock</strong> guarantees exclusive access to a critical section, even when multiple processes run on different machines.</p>
<p><strong>Redis‑based lock</strong> (the most common choice):</p>
<pre><code class="language-csharp">public class RedisLock : IDisposable
{
    private readonly IDatabase _db;
    private readonly string _key;
    private readonly TimeSpan _ttl;
    private bool _owns;

    public RedisLock(string connectionString, string key, TimeSpan ttl = default)
    {
        var mux = ConnectionMultiplexer.Connect(connectionString);
        _db = mux.GetDatabase();
        _key = key;
        _ttl = ttl == default ? TimeSpan.FromSeconds(30) : ttl;
    }

    public bool TryAcquire(TimeSpan wait)
    {
        var token = Guid.NewGuid().ToString();
        var expire = DateTime.UtcNow.Add(_ttl);
        var result = _db.LockAsync(_key, token, expire, wait);
        _owns = result;
        return _owns;
    }

    public void Release()
    {
        if (_owns)
            _db.LockRelease(_key, Guid.NewGuid().ToString());
    }

    public void Dispose() =&gt; _db.Dispose();
}
</code></pre>
<p><strong>Key points:</strong></p>
<ul>
<li><p>Use <code>SET NX PX</code> under the hood – set the key only if it doesn’t exist and give it a TTL.</p>
</li>
<li><p>The lock token (a GUID) prevents a stale lock from being released by another client.</p>
</li>
<li><p>Renew the lock inside the critical section if the job runs longer than the TTL.</p>
</li>
</ul>
<p><strong>Alternative – DB advisory lock:</strong> In SQL Server you can call <code>spgetapplock</code> <em>/</em> <code>spreleaseapplock</code>. This keeps everything in one database, avoiding an extra service.</p>
<h3>Layer 2 – Idempotent Logic</h3>
<p>Even with a lock, network partitions or process crashes can cause the lock to be released prematurely. Therefore, make the work <strong>idempotent</strong>:</p>
<ol>
<li><p><strong>Correlation ID</strong> – generate a unique ID for each message (e.g., the message’s <code>MessageId</code> or a GUID you assign when enqueuing).</p>
</li>
<li><p><strong>State table</strong> – store a record that says “processed X with correlation ID Y”.</p>
</li>
<li><p><strong>Check before act</strong> – before performing any side‑effect, query this table. If the ID already exists with a <code>Completed</code> status, skip the work.</p>
</li>
</ol>
<pre><code class="language-csharp">public class IdempotentProcessor
{
    private readonly IDbContext _ctx;

    public IdempotentProcessor(IDbContext ctx) =&gt; _ctx = ctx;

    public async Task&lt;bool&gt; AlreadyDoneAsync(string corrId)
    {
        return await _ctx.ProcessingLog.AnyAsync(p =&gt; p.CorrelationId == corrId &amp;&amp; p.Status == "Completed");
    }

    public async Task MarkDoneAsync(string corrId)
    {
        var log = new ProcessingLog
        {
            CorrelationId = corrId,
            Status = "Completed",
            ProcessedAt = DateTime.UtcNow
        };
        _ctx.ProcessingLog.Add(log);
        await _ctx.SaveChangesAsync();
    }
}
</code></pre>
<p><strong>Workflow example:</strong></p>
<pre><code class="language-csharp">var lock = new RedisLock(connectionString, $"job:{corrId}");
if (!await lock.TryAcquire(TimeSpan.FromSeconds(15))) 
    return; // another instance is handling it

if (await idempotentProcessor.AlreadyDoneAsync(corrId))
{
    await lock.Release();
    return;
}

// Critical section – safe to run
await ProcessMessageAsync(message);
await idempotentProcessor.MarkDoneAsync(corrId);
await lock.Release();
</code></pre>
<h2>4. Putting It All Together</h2>
<pre><code class="language-csharp">public async Task ProcessMessageAsync(Message msg)
{
    var corrId = msg.MessageId; // unique per message
    using var lock = new RedisLock("redis://localhost:6379", $"job:{corrId}");
    if (!await lock.TryAcquire(TimeSpan.FromSeconds(20)))
        return; // another worker has the lock

    if (await idempotentProcessor.AlreadyDoneAsync(corrId))
        return; // already finished

    // ----- real work -----
    await DoTheWorkAsync(msg);
    // ----------------

    await idempotentProcessor.MarkDoneAsync(corrId);
}
</code></pre>
<h3>Benefits</h3>
<ul>
<li><p><strong>Safety</strong> – Even if the lock expires (e.g., due to a GC pause), the idempotent check prevents duplicate side‑effects.</p>
</li>
<li><p><strong>Scalability</strong> – Multiple instances can run concurrently on <em>different</em> messages without colliding.</p>
</li>
<li><p><strong>Observability</strong> – The processing‑log table gives you a clear audit trail for debugging.</p>
</li>
</ul>
<h2>5. Practical Checklist</h2>
<table>
<thead>
<tr>
<th>✅</th>
<th>Item</th>
</tr>
</thead>
<tbody><tr>
<td>1</td>
<td>Choose a lock store (Redis is fastest; SQL Server works if you already have a transactional DB).</td>
</tr>
<tr>
<td>2</td>
<td>Define a <strong>correlation ID</strong> that is immutable for the life of the message.</td>
</tr>
<tr>
<td>3</td>
<td>Create a lightweight <strong>state table</strong> (<code>ProcessingLog</code>) with columns: <code>Id</code>, <code>CorrelationId</code>, <code>Status</code>, <code>ProcessedAt</code>.</td>
</tr>
<tr>
<td>4</td>
<td>Implement <strong>lock acquisition</strong> with TTL and renewal logic for long jobs.</td>
</tr>
<tr>
<td>5</td>
<td>Wrap the work in a <strong>try/finally</strong> block to guarantee lock release even on exceptions.</td>
</tr>
<tr>
<td>6</td>
<td>Log meaningful events (lock acquired, lock released, idempotent skip) for diagnostics.</td>
</tr>
<tr>
<td>7</td>
<td>Monitor lock contention (e.g., Redis <code>INFO LOCK</code> metrics) to detect hot keys.</td>
</tr>
</tbody></table>
<h2>6. Conclusion</h2>
<p>The race condition in scaled‑out ASP.NET Core background jobs is not a theoretical edge case—it’s a production‑ready bug that hides until traffic spikes. By <strong>pairing a distributed lock with idempotent processing</strong>, you create a robust safety net that works even when the lock momentarily fails.</p>
<p>Adopt this two‑layer pattern today: lock the critical section, then verify that the work for a given correlation ID has not already been performed. Your background jobs will become deterministic, easier to debug, and reliable at any scale.</p>
<hr />
]]></content:encoded></item><item><title><![CDATA[Why Your Playwright Tests Randomly Fail on CI and How to Fix It]]></title><description><![CDATA[Debugging Flaky Playwright Tests: Network Timing on CI
Every .NET developer who writes end-to-end tests eventually hits this wall: tests that pass on your machine, fail on CI, and seem to fail randoml]]></description><link>https://developerimranahmed.hashnode.dev/why-your-playwright-tests-randomly-fail-on-ci-and-how-to-fix-it</link><guid isPermaLink="true">https://developerimranahmed.hashnode.dev/why-your-playwright-tests-randomly-fail-on-ci-and-how-to-fix-it</guid><dc:creator><![CDATA[Imran Ahmed]]></dc:creator><pubDate>Sat, 26 Sep 2026 02:32:38 GMT</pubDate><content:encoded><![CDATA[<h1>Debugging Flaky Playwright Tests: Network Timing on CI</h1>
<p>Every .NET developer who writes end-to-end tests eventually hits this wall: tests that pass on your machine, fail on CI, and seem to fail randomly. The natural instinct is to blame the CI environment or increase timeouts. Neither approach actually fixes the problem.</p>
<p>After debugging dozens of intermittent test failures, I've found the root cause is almost always the same: tests that don't properly wait for asynchronous operations to complete.</p>
<h2>The Race Condition You're Probably Ignoring</h2>
<p>Consider this common pattern:</p>
<pre><code class="language-csharp">await page.ClickAsync("#save-button");
await page.WaitForSelectorAsync(".success-message");
</code></pre>
<p>This looks reasonable. Click the button, wait for the success message. But what if <code>#save-button</code> triggers an AJAX call that populates <code>.success-message</code>? The <code>WaitForSelectorAsync</code> might find an element that hasn't been updated yet by the network response.</p>
<p>Locally, your API responds in 50ms. The selector appears almost instantly. On CI, especially under load or in a container with network throttling, that same API might take 800ms. Your test has already moved on before the response arrives.</p>
<p>This isn't a flaky test. It's an unwaited dependency.</p>
<h2>Explicit Waits Beat Arbitrary Waits</h2>
<p>The Playwright documentation recommends explicit waits, but the difference matters:</p>
<pre><code class="language-csharp">// ❌ Arbitrary wait - guessing
await Task.Delay(2000);
await page.ClickAsync("#save-button");

// ✅ Explicit wait - knowing
await page.ClickAsync("#save-button");
await page.WaitForLoadStateAsync(LoadState.NetworkIdle);
</code></pre>
<p>The explicit approach waits exactly as long as needed. The arbitrary approach waits whether it's needed or not.</p>
<h3>Waiting for Specific Responses</h3>
<p>When you know which API call your action triggers:</p>
<pre><code class="language-csharp">await page.ClickAsync("#search-button");

// Wait for the specific response before checking results
var responsePromise = page.WaitForResponseAsync("**/api/search**");
await responsePromise;

var response = await responsePromise;
Assert.AreEqual(200, response.Status);

// Safe to interact with search results now
await page.WaitForSelectorAsync(".search-result");
</code></pre>
<h3>Waiting for Network Idle</h3>
<p>For complex interactions with multiple concurrent requests:</p>
<pre><code class="language-csharp">await page.GotoAsync("https://your-app.com/dashboard");
await page.WaitForLoadStateAsync(LoadState.NetworkIdle);

// At this point, the page has finished loading all resources
// and has no pending network requests
</code></pre>
<h3>Waiting for Selector State</h3>
<p>For elements that appear or disappear based on application logic:</p>
<pre><code class="language-csharp">await page.ClickAsync("#delete-item");

// Wait for loading indicator to appear
await page.WaitForSelectorAsync(".loading", 
    new PageWaitForSelectorOptions { State = WaitForSelectorState.Visible });

// Wait for it to disappear
await page.WaitForSelectorAsync(".loading", 
    new PageWaitForSelectorOptions { State = WaitForSelectorState.Hidden });

// Now safe to assert
Assert.IsFalse(await page.IsVisibleAsync(".item-to-delete"));
</code></pre>
<h2>Eliminating External Dependencies</h2>
<p>External APIs are a common source of CI flakiness. Third-party services have variable latency and occasional outages. In your CI pipeline, you don't want either.</p>
<p>Playwright lets you intercept and mock network requests:</p>
<pre><code class="language-csharp">await page.RouteAsync("**/api/external-service**", route =&gt;
{
    // Return mock data immediately
    route.FulfillAsync(new RouteFulfillOptions
    {
        Status = 200,
        ContentType = "application/json",
        Body = "{\"status\": \"active\", \"data\": []}"
    });
});

// Now your test runs against predictable data
await page.ClickAsync("#sync-button");
await page.WaitForSelectorAsync(".synced-data");
</code></pre>
<p>This makes your test deterministic. The external service could be down entirely and your test would still pass.</p>
<h2>Clean Browser Context Between Tests</h2>
<p>State leakage between tests is a silent flakiness generator. One test sets authentication cookies. Another test assumes a logged-out state. They run in parallel on CI and fail unpredictably.</p>
<p>Use isolated browser contexts:</p>
<pre><code class="language-csharp">private IBrowserContext _browserContext;

[SetUp]
public async Task Setup()
{
    var browser = await Playwright.Chromium.LaunchAsync();
    _browserContext = await browser.NewContextAsync();
    
    // Ensure clean state
    await _browserContext.ClearCookiesAsync();
}

[TearDown]
public async Task Teardown()
{
    await _browserContext.DisposeAsync();
}
</code></pre>
<p>For tests that need authentication, set it up explicitly rather than relying on previous test state:</p>
<pre><code class="language-csharp">[SetUp]
public async Task LoginUser()
{
    await _browserContext.ClearCookiesAsync();
    await page.GotoAsync("/login");
    await page.FillAsync("#username", "testuser");
    await page.FillAsync("#password", "testpass");
    await page.ClickAsync("#login-submit");
    await page.WaitForLoadStateAsync(LoadState.NetworkIdle);
}
</code></pre>
<h2>Diagnosing When You're Not Sure</h2>
<p>If you inherit a flaky test suite and don't know where to start, add instrumentation:</p>
<pre><code class="language-csharp">public async Task&lt;T&gt; TimedOperation&lt;T&gt;(string name, Func&lt;Task&lt;T&gt;&gt; operation)
{
    var start = DateTime.UtcNow;
    var result = await operation();
    var elapsed = (DateTime.UtcNow - start).TotalMilliseconds;
    
    Console.WriteLine($"{name} completed in {elapsed}ms");
    return result;
}

// Usage
await TimedOperation("Page load", () =&gt; page.GotoAsync(url).AsTask());
await TimedOperation("Button click", () =&gt; page.ClickAsync("#action").AsTask());
</code></pre>
<p>This data tells you where time is actually spent and which operations are the culprits.</p>
<h2>The Practical Takeaway</h2>
<p>Stop treating CI flakiness as an environmental problem. Start treating it as a diagnostic signal. When a test fails intermittently, it's usually because your test makes assumptions about timing that happen to be true locally but not in CI.</p>
<p>Fix it by:</p>
<ol>
<li><p>Using <code>waitForResponse</code> or <code>waitForLoadState</code> instead of <code>Task.Delay</code></p>
</li>
<li><p>Mocking external API calls</p>
</li>
<li><p>Isolating browser context between tests</p>
</li>
<li><p>Adding timing instrumentation to identify the real bottlenecks</p>
</li>
</ol>
<p>Your test suite becomes reliable not by waiting longer, but by waiting smarter.</p>
]]></content:encoded></item><item><title><![CDATA[The CancellationToken That Never Propagated: A Subtle ASP.NET Core Timeout Bug I Missed in Code Review]]></title><description><![CDATA[The Silent Killer of ASP.NET Core Performance: CancellationToken Propagation
I remember the first time I saw an ASP.NET Core application "hang" not because of a deadlock, but because it was doing too ]]></description><link>https://developerimranahmed.hashnode.dev/the-cancellationtoken-that-never-propagated-a-subtle-asp-net-core-timeout-bug-i-missed-in-code-review</link><guid isPermaLink="true">https://developerimranahmed.hashnode.dev/the-cancellationtoken-that-never-propagated-a-subtle-asp-net-core-timeout-bug-i-missed-in-code-review</guid><dc:creator><![CDATA[Imran Ahmed]]></dc:creator><pubDate>Fri, 25 Sep 2026 13:57:07 GMT</pubDate><content:encoded><![CDATA[<h1>The Silent Killer of ASP.NET Core Performance: CancellationToken Propagation</h1>
<p>I remember the first time I saw an ASP.NET Core application "hang" not because of a deadlock, but because it was doing too much work. It wasn't crashing. It was just... slow. And the reason was a subtle bug in how we handled background tasks.</p>
<p>This is a story about a <code>CancellationToken</code> that never propagated, and why that single omission can turn a performant API into a resource-hungry beast under load.</p>
<h2>The Symptom</h2>
<p>We were using a common pattern: fast response, background processing.</p>
<pre><code class="language-csharp">public async Task&lt;IActionResult&gt; SendEmail(string email)
{
    _ = SendEmailAsync(email); // Fire and forget
    return Ok("Email queued");
}
</code></pre>
<p>This seemed fine. The user gets a 200 OK immediately. The email gets sent in the background. But during a marketing campaign, where thousands of emails were queued, the API started responding slowly to <em>other</em> requests.</p>
<h2>The Root Cause</h2>
<p>We assumed that because <code>SendEmailAsync</code> was <code>async</code>, it wouldn't block the main thread. And it didn't. But here’s the catch: <code>async</code> doesn’t mean "short-lived."</p>
<p><code>SendEmailAsync</code> was calling a third-party SMTP server that was sometimes slow. When the SMTP server was down, it would retry 5 times with a 10-second delay between each. That’s a 50-second window of waiting.</p>
<p>Since we never passed a <code>CancellationToken</code>, if the application decided to shut down (or if the request context that spawned it was cancelled), the background task kept running. It kept waiting. It kept occupying a thread pool thread (or at least an async context slot).</p>
<p>When we had 5,000 pending emails, we had 5,000 tasks waiting on the slow SMTP server. The thread pool saturation was inevitable.</p>
<h2>The Solution: Cooperative Cancellation</h2>
<p>The fix required changing the architecture slightly. We couldn't just "fix" the method; we had to change the contract.</p>
<ol>
<li><p><strong>Accept a Token:</strong> <code>SendEmailAsync</code> now takes a <code>CancellationToken</code>.</p>
</li>
<li><p><strong>Provide a Token:</strong> The caller provides a token. For request-scoped operations, we used <code>HttpContext.RequestAborted</code>. For global shutdown, we hooked into the <code>ApplicationStopping</code> event.</p>
</li>
</ol>
<pre><code class="language-csharp">public async Task&lt;IActionResult&gt; SendEmail(string email, CancellationToken cancellationToken)
{
    // Pass the token to the background task
    _ = SendEmailAsync(email, cancellationToken);
    return Ok("Email queued");
}

public async Task SendEmailAsync(string email, CancellationToken token)
{
    for (int i = 0; i &lt; 5; i++)
    {
        try
        {
            await smtpClient.SendMailAsync(email, token);
            break;
        }
        catch (OperationCanceledException)
        {
            // The app is shutting down or the client disconnected.
            // Stop retrying.
            return;
        }
        catch (Exception)
        {
            // Wait before retrying, but check if we've been cancelled
            await Task.Delay(TimeSpan.FromSeconds(10), token);
        }
    }
}
</code></pre>
<h2>Why Timeouts Aren't Enough</h2>
<p>Developers often confuse <code>Timeouts</code> with <code>Cancellation</code>.</p>
<ul>
<li><p><strong>Timeout:</strong> "Give up after 30 seconds." This releases the <em>caller</em> from waiting.</p>
</li>
<li><p><strong>Cancellation:</strong> "Stop working now." This releases the <em>system</em> from performing work.</p>
</li>
</ul>
<p>If you only have a timeout, the background work might still be running in memory. If you have cancellation, you can proactively kill that work. Under load, this difference is the difference between a green dashboard and a red one.</p>
<h2>Key Takeaways</h2>
<ul>
<li><p><strong>Fire-and-Forget is Dangerous:</strong> Never fire-and-forget a task that does significant work without a cancellation path.</p>
</li>
<li><p><strong>Check the Token:</strong> Inside long loops or delays, always check <code>token.IsCancellationRequested</code>.</p>
</li>
<li><p><strong>Use Standard Tokens:</strong> Use <code>RequestAborted</code> for web requests. Use <code>Host.StopAsync</code> token for background services.</p>
</li>
<li><p><strong>Review for "Silent Work":</strong> In code reviews, look for <code>Task.Run</code> or detached async calls. Ask: "How does this stop?"</p>
</li>
</ul>
<p>If you want to learn more about advanced .NET concurrency, check out the resources linked below.</p>
<p>#CSharp #DotNet #Performance #Backend</p>
]]></content:encoded></item></channel></rss>