• About Blog

    What's Blog?

    A blog is a discussion or informational website published on the World Wide Web consisting of discrete, often informal diary-style text entries or posts.

  • About Cauvery Calling

    Cauvery Calling. Action Now!

    Cauvery Calling is a first of its kind campaign, setting the standard for how India’s rivers – the country’s lifelines – can be revitalized.

  • About Quinbay Publications

    Quinbay Publication

    We follow our passion for digital innovation. Our high performing team comprising of talented and committed engineers are building the future of business tech.

Showing posts with label Spring Boot. Show all posts
Showing posts with label Spring Boot. Show all posts

Wednesday, June 18, 2025

Let's Build a Storyteller with Spring AI

Photo Courtesy DevDocsMaster

Remember those childhood summer holidays? After a long day of cricket in the neighbourhood lane, we’d all gather around, and there was always someone—a grandparent, an uncle, or an older cousin—who was the master storyteller.

I still remember my Mom, she could spin up the most fascinating tales out of thin air. Stories of a clever fox who outsmarted a lion, a tiny sparrow on a big adventure, or a king who learned a lesson from a poor farmer. We’d listen, completely captivated, our imaginations painting vivid pictures. Those simple stories were a magical part of growing up.

Now, as developers, what if we could bring a slice of that magic into our digital world? How about we build our very own storyteller? An application where you give it a tiny spark of an idea—say, "a curious robot who discovers desi chai"—and it instantly writes a wonderful short story for you.

Sounds like a fun project, right? But my mind immediately jumps to the challenges. Figuring out complex AI libraries, handling API calls, all that backend hassle… it seems like it would take all the fun out of it. How can we build something so creative using our solid, reliable Java and Spring Boot?

Well, this is where the story gets really interesting for us. It turns out, the brilliant minds at Spring have already thought about this. And their answer is Spring AI.

So, What’s All This Hungama About Spring AI?

Think of Spring AI as a friendly bridge. On one side, you have your solid, dependable Spring Boot application. On the other, you have the incredible power of AI models like OpenAI's GPT, Google's Gemini, and others. Spring AI connects these two worlds so seamlessly that you'll wonder why you ever thought AI was difficult.

In simple terms, it takes away all the boilerplate code and complex configurations. You don't have to manually handle HTTP requests to AI services or parse messy JSON responses. Spring AI gives you a clean, straightforward way to talk to AI, just like you would talk to any other service in your Spring application.

Let's Build Something! Your First AI-Powered Spring Boot App

Enough talk, let’s get our hands dirty. Let's build our little "Story Generator." You need to give a simple idea, and it cooks up a short story for you.

We'll be building this faster than it takes to get your food delivery on a Friday night.

Step 1: The Foundation - Setting Up Your Project

First things first, we need a basic Spring Boot project. The easiest way is to use the Spring Initializr. It’s our go-to starting point for any new Spring project.

  1. Head over to start.spring.io.
  2. Choose Maven as the project type and Java as the language.
  3. Select a recent stable version of Spring Boot (3.2.x or higher is good).
  4. Give your project a name, something like ai-story-generator.
  5. Now, for the important part – the dependencies. Add the following:
    • Spring Web: Because we want to create a REST endpoint.
    • Spring Boot Actuator: Good practice to monitor our app.
    • OpenAI: This is the Spring AI magic wand we need. Just type "OpenAI" and add the dependency.

Once you’re done, click "Generate". A zip file will be downloaded. Unzip it and open the project in your favourite IDE (IntelliJ or VS Code, your choice!).

Step 2: The Secret Ingredient - Your API Key

To talk to an AI model like OpenAI's, you need an API key. It's like a secret password.

  1. Go to the OpenAI Platform and create an account.
  2. Navigate to the API Keys section and create a new secret key.
  3. Important: Copy this key immediately and save it somewhere safe. You won’t be able to see it again!

Now, open the src/main/resources/application.properties file in your project and add this line:

spring.ai.openai.api-key=YOUR_OPENAI_API_KEY_HERE

Step 3: Writing the Code - Where the Magic Happens

This is the best part. You'll be surprised at how little code we need to write.

Let's create a simple REST controller. Create a new Java class called StoryController.java.

package com.bhargav.ai.storygenerator;

import org.springframework.ai.chat.client.ChatClient;
import org.springframework.web.bind.annotation.GetMapping;
import org.springframework.web.bind.annotation.RequestParam;
import org.springframework.web.bind.annotation.RestController;

@RestController
public class StoryController {

    private final ChatClient chatClient;

    public StoryController(ChatClient.Builder chatClientBuilder) {
        this.chatClient = chatClientBuilder.build();
    }

    @GetMapping("/story")
    public String generateStory(@RequestParam(value = "topic",
            defaultValue = "a curious robot who discovered desi chai") String topic) {
        return this.chatClient.prompt()
                .user("Tell me a short story about " + topic)
                .call()
                .content();
    }
}

Let's break down this simple code, shall we?

  • @RestController: This tells Spring that this class will handle web requests.
  • private final ChatClient chatClient;: This is the hero of our story! The ChatClient is a part of Spring AI that makes talking to the AI model incredibly easy. We inject it using the constructor. Spring Boot automatically configures it for us because we added the OpenAI dependency and the API key. No manual setup needed. Kitna aasan hai! (How easy is that!)
  • @GetMapping("/story"): This creates a web endpoint. You can access it at http://localhost:8080/story.
  • The generateStory method is where the action is.
  • chatClient.prompt(): We start building our request to the AI.
  • .user("Tell me a short story about " + topic): We are telling the AI what to do. This is our "prompt." We take a topic from the user's request.
  • .call(): This sends our request to the AI model.
  • .content(): This gets the text response back from the AI.

And that’s it! We’re done. Seriously.

Step 4: Run the Application!

Now, just run your Spring Boot application from your IDE. Once it starts up, open your web browser and go to:

http://localhost:8080/story

You should see a short story about a curious robot discovering chai.

Want to try another topic? Just add a topic parameter to the URL:

http://localhost:8080/story?topic=a cat who wanted to be a software engineer in Bengaluru

And watch as the AI instantly generates a new story for you.

What Did We Just Do?

Think about it. In just a few minutes, with a handful of dependencies and less than 20 lines of Java code, we built an AI-powered application. We didn't have to wrestle with HTTP clients, authentication headers, or complex JSON.

We just told Spring AI what we wanted, and it did the needful.

This is just the tip of the iceberg. Spring AI allows you to get structured output (like JSON objects), connect to your own data, and much more. It makes AI a first-class citizen in the Spring ecosystem.

So, the next time you feel that spark of a creative idea, don't think it's out of reach for a Java developer. With Spring AI in your toolkit, you're more than ready to build your own magic. Happy coding!


Thursday, January 21, 2021

Rate Limiter Implementation — Sliding Log Algorithm

Sliding Log Image

API Rate Limiting

Rate limiting is a strategy to limit the access to APIs. It restricts the number of API calls that a client can make within any given timeframe. This helps to defend the API against abuse, both unintentional and malicious scripts.

Rate limits are often applied to an API by tracking the IP address, API keys or access tokens, etc. As an API developers, we can choose to respond in several different ways when a client reaches the limit.

  • Queueing the request until the remaining time period has elapsed.
  • Allowing the request immediately but charging extra for this request.
  • Most common one is rejecting the request (HTTP 429 Too Many Requests)

Sliding Log Algorithm

Sliding Log rate limiting involves tracking a time stamped log for each consumer request. These logs are usually stored in a hash set or table that is sorted by time. Logs with timestamps beyond a threshold are discarded. When a new request comes in, we calculate the sum of logs to determine the request rate. If the request would exceed the threshold rate, then it is held.

The advantage of this algorithm is that it does not suffer from the boundary conditions of fixed windows. The rate limit will be enforced precisely and because the sliding log is tracked for each consumer, you don’t have the rush effect that challenges fixed windows. However, it can be very expensive to store an unlimited number of logs for every request. It’s also expensive to compute because each request requires calculating a summation over the consumers prior requests, potentially across a cluster of servers. As a result, it does not scale well to handle large bursts of traffic or denial of service attacks.

Please refer to the Understanding Rate Limiting Algorithms blog where the Sliding Log and other algorithms have been explained in detail.

Building a Springboot Application with API Rate Limiter

Create a new spring boot application from Spring Initializr with dependency on spring web module.

Unzip the downloaded project and import to your IDE. We are going to implement a simple calculator REST APIs that can do operations like add and subtract.

@RestController
@RequestMapping(value = "/api/calculator")
public class CalculatorController {
    @GetMapping(value = "/add")
    public ResponseEntity<Calculator> add(@RequestParam int left, @RequestParam int right) {
        return ResponseEntity.ok(Calculator.builder()
                .operation("add").answer(left + right).build());
    }
    @GetMapping(value = "/subtract")
    public ResponseEntity<Calculator> subtract(@RequestParam int left, @RequestParam int right) {
        return ResponseEntity.ok(Calculator.builder()
                .operation("subtract").answer(left - right).build());
    }
}

Let’s ensure that our above APIs are up and running as expected. You can use the cURL or PostMan to make an API call.

curl -X GET -H "Content-Type: application/json" 'http://localhost:9090/api/calculator/add?left=20&right=30'{"operation":"add","answer":50}

Now that we have APIs ready to consume, next let’s introduce some subscription plans with rate limits. Let’s assume that we have the following subscription plans for our clients:
Free Subscription allows 2 requests per 60 seconds.
Basic Subscription allows 10 requests per 60 seconds.
Professional Subscription allows 20 requests per 60 seconds.

Each API client gets a unique API key that they must send along with each request. This would help us identify the client and subscription plan linked.

public enum SubscriptionPlan {

    SUBSCRIPTION_FREE(2, 60),
    SUBSCRIPTION_BASIC(10, 60),
    SUBSCRIPTION_PROFESSIONAL(20, 60);

    private final int requestLimit;
    private final int windowTime;

    SubscriptionPlan(int requestLimit, int windowTime) {
        this.requestLimit = requestLimit;
        this.windowTime = windowTime;
    }

    public int getRequestLimit() {
        return this.requestLimit;
    }

    public int getWindowTime() {
        return this.windowTime;
    }

}

Next we create a subscription service which will store the references for each of the API client in a memory.

@Service
public class SubscriptionService {

    private final Map<String, UserRequestData>
            subscriptionCacheMap = new ConcurrentHashMap<>();

    public UserRequestData resolveSubscribedUserData(String subscriptionKey) {
        return subscriptionCacheMap.computeIfAbsent(
                subscriptionKey, this::resolveUser);
    }

    private UserRequestData resolveUser(String subscriptionKey) {
        if (subscriptionCacheMap.containsKey(subscriptionKey)) {
            return subscriptionCacheMap.get(subscriptionKey);
        }
        return buildUserLog(
                resolveSubscriptionPlanByKey(subscriptionKey));
    }

    private UserRequestData buildUserLog(SubscriptionPlan subscriptionPlan) {
        return new UserRequestData(
                subscriptionPlan.getRequestLimit(), 
                subscriptionPlan.getWindowTime());
    }

    private SubscriptionPlan resolveSubscriptionPlanByKey(String subscriptionKey) {
        if (subscriptionKey.startsWith("PS1129-")) {
            return SubscriptionPlan.SUBSCRIPTION_PROFESSIONAL;
        } else if (subscriptionKey.startsWith("BS1129-")) {
            return SubscriptionPlan.SUBSCRIPTION_BASIC;
        }

        return SubscriptionPlan.SUBSCRIPTION_FREE;
    }

}

Let’s understand the implementation. The API client sends an API key with the X-Subscription-Key request header. We use the SubscriptionService to get the user reference for the API key and check whether the request is allowed or not with the help of methods.

In order to enhance the client experience of the API, we will add the following additional response headers to send information about the rate limit.
  • X-Rate-Limit-Remaining — number of tokens remaining in the current time window.
  • X-Rate-Limit-Retry-After-Seconds — remaining time in seconds until the bucket is refilled with new tokens.
We can call UserRequestData methods getRequestWaitTime and getRemainingRequests, to get the count of the remaining requests and the time remaining until the next sliding log respectively. The implementation provided in this class is self explanatory and easy to understand the same.

public class UserRequestData {

    private int requestLimit;
    private int windowTimeInSec;
    private Queue<Long> requestTimeStamps;

    public UserRequestData(
            int requestLimit, int windowTimeInSec) {
        this.requestLimit = requestLimit;
        this.windowTimeInSec = windowTimeInSec;
        this.requestTimeStamps =
                new ConcurrentLinkedDeque<Long>();
    }

    public int getRemainingRequests() {
        return requestLimit - requestTimeStamps.size();
    }

    public int getRequestWaitTime() {
        long currentTimeStamp =
                System.currentTimeMillis() / 1000;
        int initialElapsedTime =
                (int) (currentTimeStamp - requestTimeStamps.peek());
        return initialElapsedTime > windowTimeInSec
                ? 0 : windowTimeInSec - initialElapsedTime;
    }

    public boolean isServiceCallAllowed() {
        long currentTimeStamp =
                System.currentTimeMillis() / 1000;
        evictOlderRequestTimeStamps(currentTimeStamp);

        if (requestTimeStamps.size() >= this.requestLimit) {
            return false;
        }

        requestTimeStamps.add(currentTimeStamp);
        return true;
    }

    public void evictOlderRequestTimeStamps(long currentTimeStamp) {
        while (requestTimeStamps.size() != 0 &&
                (currentTimeStamp - requestTimeStamps.peek() 
                        > windowTimeInSec)) {
            requestTimeStamps.remove();
        }
    }

}

Here is the implementation of the Interceptor to validate the request with rate limiter to see whether we accept or reject the request.

@Component
public class RateLimiterInterceptor implements HandlerInterceptor {

    private static final String
            HEADER_SUBSCRIPTION_KEY = "X-Subscription-Key";
    private static final String
            HEADER_LIMIT_REMAINING = "X-Rate-Limit-Remaining";
    private static final String
            HEADER_RETRY_AFTER = "X-Rate-Limit-Retry-After-Seconds";
    private static final String
            SUBSCRIPTION_QUOTA_EXHAUSTED =
            "You've exhausted your API Request Quota. " +
            "Please upgrade your subscription plan.";

    @Autowired
    private SubscriptionService subscriptionService;

    @Override
    public boolean preHandle(HttpServletRequest request,
                             HttpServletResponse response,
                             Object handler) throws Exception {
        String subscriptionKey =
                request.getHeader(HEADER_SUBSCRIPTION_KEY);
        if (StringUtils.isEmpty(subscriptionKey)) {
            response.sendError(HttpStatus.BAD_REQUEST.value(),
                    "Missing Request Header: " +
                       HEADER_SUBSCRIPTION_KEY);
            return false;
        }

        UserRequestData userRequestData = subscriptionService
                        .resolveSubscribedUserData(subscriptionKey);
        if (!userRequestData.isServiceCallAllowed()) {
            int waitTime = userRequestData.getRequestWaitTime();
            response.addHeader(HEADER_RETRY_AFTER,
                    String.valueOf(waitTime));

            response.setContentType(
                    MediaType.APPLICATION_JSON_VALUE);
            response.sendError(
                    HttpStatus.TOO_MANY_REQUESTS.value(),
                    SUBSCRIPTION_QUOTA_EXHAUSTED);
            return false;
        }

        response.addHeader(HEADER_LIMIT_REMAINING,
                String.valueOf(
                    userRequestData.getRemainingRequests()));
        return true;
    }
}

Finally, let’s add the interceptor to the InterceptorRegistry of Springboot so that the RateLimitInterceptor intercepts each request to our calculator API endpoints.

@SpringBootApplication
public class SlidingWindowApplication implements WebMvcConfigurer {

    @Autowired
    @Lazy
    private RateLimiterInterceptor interceptor;

    public void addInterceptors(InterceptorRegistry registry) {
        registry.addInterceptor(interceptor)
                .addPathPatterns("/api/calculator/**");
    }

    public static void main(String[] args) {
        SpringApplication.run(SlidingWindowApplication.class, args);
    }

}

Let invoke calculator API to see the behaviour.

curl -X GET 'http://localhost:9090/api/calculator/add?left=20&right=30'
{"timestamp":"2021-01-03T12:56:20.047+0000","status":400,"error":"Bad Request","message":"Missing Request Header: X-Subscription-Key","path":"/api/calculator/add"}

The client has to send the API key within the http header otherwise the interceptor will not process the request. Let’s add the API key to the header and make the call.

curl -v -X GET -H "X-subscription-key:A1129-12" 'http://localhost:9090/api/calculator/subtract?left=20&right=30'
* Connected to localhost (::1) port 9090 (#0)
> GET /api/calculator/subtract?left=20&right=30 HTTP/1.1
> Host: localhost:9090
> User-Agent: curl/7.64.1
> Accept: */*
> X-subscription-key:A1129-12
>
< HTTP/1.1 200
< X-Rate-Limit-Remaining: 1
< Content-Type: application/json
< Transfer-Encoding: chunked
< Date: Sun, 03 Jan 2021 12:57:09 GMT
<
* Connection #0 to host localhost left intact
{"operation":"subtract","answer":-10}
* Closing connection 0

You can see the API key is added in the header, the API responds to our request and also it has added response header which shows how many rate is remaining for the API key.

Let’s make 2 more calls then we should see that we exhausted our rate for the free plan and returns 429 as response.

curl -v -X GET -H "X-subscription-key:A1129-12" 'http://localhost:9090/api/calculator/subtract?left=20&right=30'
* Connected to localhost (::1) port 9090 (#0)
> GET /api/calculator/subtract?left=20&right=30 HTTP/1.1
> Host: localhost:9090
> User-Agent: curl/7.64.1
> Accept: */*
> X-subscription-key:A1129-12
>
< HTTP/1.1 429
< X-Rate-Limit-Retry-After-Seconds: 24
< Content-Type: application/json
< Transfer-Encoding: chunked
< Date: Sun, 03 Jan 2021 12:58:58 GMT
<
* Connection #0 to host localhost left intact
{"timestamp":"2021-01-03T12:58:58.176+0000","status":429,"error":"Too Many Requests","message":"You've exhausted your API Request Quota. Please upgrade your subscription plan.","path":"/api/calculator/subtract"}
* Closing connection 0

It looks like we have successfully implemented the rate limiter using the Sliding Log algorithm. We can keep adding endpoints and the interceptor would apply the rate limit for each request.

As usual, the source code for the above spring boot implementation is available over on GitHub.

Monday, January 4, 2021

Rate Limiter Implementation — Token Bucket Algorithm

Token Bucket Image

API Rate Limiting

Rate limiting is a strategy to limit the access to APIs. It restricts the number of API calls that a client can make within any given timeframe. This helps to defend the API against abuse, both unintentional and malicious scripts.

Rate limits are often applied to an API by tracking the IP address, API keys or access tokens, etc. As an API developers, we can choose to respond in several different ways when a client reaches the limit.

  • Queueing the request until the remaining time period has elapsed.
  • Allowing the request immediately but charging extra for this request.
  • Most common one is rejecting the request (HTTP 429 Too Many Requests)

Token Bucket Algorithm

Assume that we have a bucket, the capacity is defined as the number of tokens that it can hold. Whenever a consumer wants to access an API endpoint, it must get a token from the bucket. Token is removed from the bucket if it’s available and accept the request. If the token is not available then the server rejects the request.

As requests are consuming tokens, we also need to refill them at some fixed rate and time, such that we never exceed the capacity of the bucket. Let’s consider an API that has a rate limit of 100 requests per minute. We can create a bucket with a capacity of 100, and a refill rate of 100 tokens per minute.

Please refer to the Understanding Rate Limiting Algorithms blog where the Token Bucket and other algorithms have been explained in detail.

Building a Springboot Application with API Rate Limiter

Create a new spring boot application from Spring Initializr with dependency on spring web module.

Unzip the downloaded project and import to your IDE. Let’s begin by adding the bucket4j dependency to our pom.xml

<dependency>
    <groupId>com.github.vladimir-bukhtoyarov</groupId>
    <artifactId>bucket4j-core</artifactId>
    <version>4.10.0</version>
</dependency>

We are going to implement a simple calculator REST APIs that can do operations like add and subtract.

@RestController
@RequestMapping(value = "/api/calculator")
public class CalculatorController {

    @GetMapping(value = "/add")
    public ResponseEntity<Calculator> add(@RequestParam int left, @RequestParam int right) {
        return ResponseEntity.ok(Calculator.builder()
                .operation("add").answer(left + right).build());
    }

    @GetMapping(value = "/subtract")
    public ResponseEntity<Calculator> subtract(@RequestParam int left, @RequestParam int right) {
        return ResponseEntity.ok(Calculator.builder()
                .operation("subtract").answer(left - right).build());
    }

}

Let’s ensure that our above APIs are up and running as expected. You can use the cURL or PostMan to make an API call.

curl -X GET -H "Content-Type: application/json" 'http://localhost:9090/api/calculator/add?left=20&right=30'
{"operation":"add","answer":50}

Now that we have APIs ready to consume, next let’s introduce some subscription plans with rate limits. Let’s assume that we have the following subscription plans for our clients:
  • Free Subscription allows 2 requests per 60 seconds.
  • Basic Subscription allows 10 requests per 60 seconds.
  • Professional Subscription allows 20 requests per 60 seconds.

Each API client gets a unique API key that they must send along with each request. This would help us identify the client and subscription plan linked.

public enum SubscriptionPlan {

    SUBSCRIPTION_FREE(2),
    SUBSCRIPTION_BASIC(10),
    SUBSCRIPTION_PROFESSIONAL(20);

    private int bucketLimit;

    private SubscriptionPlan(int bucketLimit) {
        this.bucketLimit = bucketLimit;
    }

    public int getBucketLimit() {
        return this.bucketLimit;
    }

    public Bandwidth getBandwidth() {
        return Bandwidth.classic(bucketLimit,
                Refill.intervally(bucketLimit,
                        Duration.ofMinutes(1)));
    }

}

Next we create a subscription service which will store the bucket reference for each of the API client in a memory.

@Service
public class SubscriptionService {

    private final Map<String, Bucket>
            subscriptionCacheMap = new ConcurrentHashMap<>();

    public Bucket resolveBucket(String subscriptionKey) {
        return subscriptionCacheMap.computeIfAbsent(
                subscriptionKey, this::getSubscriptionBucket);
    }

    private Bucket getSubscriptionBucket(String subscriptionKey) {
        return buildBucket(
                resolveSubscriptionPlanByKey(subscriptionKey)
                        .getBandwidth());
    }

    private Bucket buildBucket(Bandwidth limit) {
        return Bucket4j.builder().addLimit(limit).build();
    }

    private SubscriptionPlan resolveSubscriptionPlanByKey(
            String subscriptionKey) {
        if (subscriptionKey.startsWith("PS1129-")) {
            return SubscriptionPlan.SUBSCRIPTION_PROFESSIONAL;
        } else if (subscriptionKey.startsWith("BS1129-")) {
            return SubscriptionPlan.SUBSCRIPTION_BASIC;
        }

        return SubscriptionPlan.SUBSCRIPTION_FREE;
    }
}

Let’s understand the implementation. The API client sends an API key with the X-Subscription-Key request header. We use the SubscriptionService to get the bucket for this API key and check whether the request is allowed by consuming a token from the bucket.

In order to enhance the client experience of the API, we will add the following additional response headers to send information about the rate limit.
  • X-Rate-Limit-Remaining - number of tokens remaining in the current time window.
  • X-Rate-Limit-Retry-After-Seconds - remaining time in seconds until the bucket is refilled with new tokens.
We can call ConsumptionProbe methods getRemainingTokens and getNanosToWaitForRefill, to get the count of the remaining tokens in the bucket and the time remaining until the next refill, respectively. The getNanosToWaitForRefill method returns 0 if we are able to consume the token successfully.

Let’s create a RateLimitInterceptor and implement the rate limit code in the preHandle method instead of writing in every API method as we will have cleaner implementation.

@Component
public class RateLimiterInterceptor implements HandlerInterceptor {

    private static final String
            HEADER_SUBSCRIPTION_KEY = "X-Subscription-Key";
    private static final String
            HEADER_LIMIT_REMAINING = "X-Rate-Limit-Remaining";
    private static final String
            HEADER_RETRY_AFTER = "X-Rate-Limit-Retry-After-Seconds";
    private static final String
            SUBSCRIPTION_QUOTA_EXHAUSTED =
            "You've exhausted your API Request Quota. " +
            "Please upgrade your subscription plan.";

    @Autowired
    private SubscriptionService subscriptionService;

    @Override
    public boolean preHandle(HttpServletRequest request,
                             HttpServletResponse response,
                             Object handler) throws Exception {
        String subscriptionKey =
                request.getHeader(HEADER_SUBSCRIPTION_KEY);
        if (StringUtils.isEmpty(subscriptionKey)) {
            response.sendError(HttpStatus.BAD_REQUEST.value(),
                    "Missing Request Header: " +
                        HEADER_SUBSCRIPTION_KEY);
            return false;
        }

        Bucket tokenBucket = subscriptionService
                .resolveBucket(subscriptionKey);
        ConsumptionProbe consumptionProbe =
                tokenBucket.tryConsumeAndReturnRemaining(1);
        if (!consumptionProbe.isConsumed()) {
            long waitTime =
                    consumptionProbe.getNanosToWaitForRefill()
                            / 1_000_000_000;
            response.addHeader(HEADER_RETRY_AFTER,
                    String.valueOf(waitTime));

            response.setContentType(
                    MediaType.APPLICATION_JSON_VALUE);
            response.sendError(
                    HttpStatus.TOO_MANY_REQUESTS.value(),
                    SUBSCRIPTION_QUOTA_EXHAUSTED);
            return false;
        }

        response.addHeader(HEADER_LIMIT_REMAINING,
                String.valueOf(
                    consumptionProbe.getRemainingTokens()));
        
        return true;
    }
}

Finally, let’s add the interceptor to the InterceptorRegistry of Springboot so that the RateLimitInterceptor intercepts each request to our calculator API endpoints.

@SpringBootApplication
public class TokenBucketApplication implements WebMvcConfigurer {

   @Autowired
   @Lazy
   private RateLimiterInterceptor interceptor;

   public void addInterceptors(InterceptorRegistry registry) {
      registry.addInterceptor(interceptor)
            .addPathPatterns("/api/calculator/**");
   }

   public static void main(String[] args) {
      SpringApplication.run(TokenBucketApplication.class, args);
   }

}

Let invoke calculator API to see the behaviour.

curl -X GET -H "Content-Type: application/json" 'http://localhost:9090/api/calculator/add?left=20&right=30'
{"timestamp":"2020-12-25T12:43:43.239+0000","status":400,"error":"Bad Request","message":"Missing Request Header: X-Subscription-Key","path":"/api/calculator/add"}

The client has to send the API key within the http header otherwise the interceptor will not process the request. Let’s add the API key to the header and make the call.

curl -v -X GET -H "Content-Type: application/json" -H "X-subscription-key:A1129-12" 'http://localhost:9090/api/calculator/add?left=20&right=30'
* Connected to localhost (::1) port 9090 (#0)
> GET /api/calculator/add?left=20&right=30 HTTP/1.1
> Host: localhost:9090
> User-Agent: curl/7.64.1
> Accept: */*
> Content-Type: application/json
> X-subscription-key:A1129-12
>
< HTTP/1.1 200
< X-Rate-Limit-Remaining: 1
< Content-Type: application/json
< Transfer-Encoding: chunked
< Date: Fri, 25 Dec 2020 12:46:06 GMT
<
* Connection #0 to host localhost left intact
{"operation":"add","answer":50}
* Closing connection 0

You can see the API key is added in the header, the API responds to our request and also it has added response header which shows how many rate is remaining for the API key.

Let’s make 2 more calls then we should see that we exhausted our rate for the free plan and returns 429 as response.

curl -v -X GET -H "Content-Type: application/json" -H "X-subscription-key:A1129-12" 'http://localhost:9090/api/calculator/add?left=20&right=30'
* Connected to localhost (::1) port 9090 (#0)
> GET /api/calculator/add?left=20&right=30 HTTP/1.1
> Host: localhost:9090
> User-Agent: curl/7.64.1
> Accept: */*
> Content-Type: application/json
> X-subscription-key:A1129-12
>
< HTTP/1.1 429
< X-Rate-Limit-Retry-After-Seconds: 51
< Content-Type: application/json
< Transfer-Encoding: chunked
< Date: Fri, 25 Dec 2020 12:49:11 GMT
<
* Connection #0 to host localhost left intact
{"timestamp":"2020-12-25T12:49:11.358+0000","status":429,"error":"Too Many Requests","message":"You've exhausted your API Request Quota. Please upgrade your subscription plan.","path":"/api/calculator/add"}
* Closing connection 0

It looks like we have successfully implemented the rate limiter using the Token Bucket algorithm. We can keep adding endpoints and the interceptor would apply the rate limit for each request.

As usual, the source code for the above spring boot implementation is available over on GitHub.

Featured Post

Your AI Sidekick: How Claude took over Pritee’s Repetitive tasks

  It was a classic Wednesday morning in our Bengaluru office . Pritee, one of our sharpest Project Managers, had just stepped out of a stake...