# ExamBuddy: First Steps in AI Integration

A while ago, I created an application to help my wife prepare for exams. It allows users to create exams with sets of questions and answers, take those exams, and save the results.

Recently, I decided to learn something new by integrating an AI-powered feature into the application. I wanted to find out how difficult it would be to add this type of functionality to an existing system. It would also be a good test of the application’s architecture, revealing what needs to be considered and where potential bottlenecks might arise.

My goal was to introduce a feature that allows users to ask an LLM to review a question and its answers while editing them. The model would assess their correctness, quality, completeness, and wording, and the user could then either apply or reject its suggestions.

![](https://cdn.hashnode.com/uploads/covers/64d3deaccda572f1d04ef00a/614e3435-1055-4bf4-b0c6-96a6ef4596e9.png align="center")

## **What I have learned?**

*   From a coding perspective, the integration was quite straightforward.
    
*   It is sometimes worth revisiting the application’s architecture and improving it where possible. I thought my application was close to following Vertical Slice Architecture, but I was wrong.
    
*   Writing the system prompt for reviewing questions was challenging and required several iterations.
    
*   Selecting the right model was not as easy as I had expected.
    
*   Running Ollama on my local network turned out to be a great idea [(see this blog post).](https://mateusz-czernek.pl/make-ai-models-available-in-your-home-network-second-life-for-a-gaming-laptop)
    

## **Exam buddy repository**

**Complete code of the application available** [**here**](https://bitbucket.org/mat-czernek/exambuddy/src/main/)

## **Code perspective**

Required packages

```plaintext
Microsoft.Extensions.AI
OllamaSharp - or any other package that supports your AI provider
```

Services registration

```csharp
// Ollama API client - came from OllamaSharp
services.AddSingleton(sp =>
{
    var settings = sp.GetRequiredService<IOptions<AiOptions>>().Value;
    return new OllamaApiClient(new Uri(settings.Endpoint), settings.ChatModel);
});

// Registering Ollama API client as IChatClient that comes from Microsoft.Extensions.Ai
services.AddSingleton<IChatClient>(sp => sp.GetRequiredService<OllamaApiClient>());
services.AddSingleton(sp =>
{
    var settings = sp.GetRequiredService<IOptions<AiOptions>>().Value;
    return new AiRequestGate(settings.MaxConcurrentRequests);
});
```

, and that’s basically all there is to setting up the framework. If needed, I can switch to a different AI provider in the future. The only requirement is that it supports the IChatClient contract, which is currently the standard abstraction.

Next, we need a service that communicates with the chat model:

```csharp
services.AddScoped<IQuestionReviewService, QuestionReviewService>();
```

The service then simply calls a method provided by IChatClient:

```csharp
await chatClient.GetResponseAsync<CreateQuestionReview.ModelResponse>(
    messages,
    ModelResponseJsonOptions,
    chatOptions,
    useJsonSchemaResponseFormat: true,
    cancellationToken: timeoutToken);
```

The rest of the code is responsible for providing the system prompt through ModelChatSystemPromptProvider, parsing the model’s response, and transforming it into the payload expected by the UI.

As mentioned earlier, designing the system prompt was challenging. The same model would sometimes return data in different formats for the same prompt and question, while different models had their own interpretations of the expected response structure. I eventually solved this by defining the required response format directly in the system prompt:

```json
{
  "qualityLevel": "good",
  "issues": [
    {
      "category": "clarity",
      "message": "A concise actionable issue in responseLanguage."
    }
  ],
  "answerAssessments": [
    {
      "answerIndex": 0,
      "shouldBeCorrect": true,
      "explanation": "A concise explanation in responseLanguage."
    }
  ],
  "suggestedQuestionText": null,
  "suggestedAnswers": null,
  "summary": "A concise overall assessment in responseLanguage.",
  "isRefusal": false
}

When providing suggestedAnswers instead of null, use this shape:
[
  {
    "text": "Suggested answer text",
    "isCorrect": true
  }
]
```

I ended up with a system prompt tightly coupled to the expected response payload, but that was an acceptable trade-off. The response contract is covered by an integration test that runs against a real model through Ollama. This ensures that any mismatch caused by future changes to the response structure is detected, allowing me to update the prompt accordingly.

See the test here: [https://bitbucket.org/mat-czernek/exambuddy/src/main/ExamBuddy.Api.Tests/Features/QuestionReviews/QuestionReviewOllamaLocalModel.cs](https://bitbucket.org/mat-czernek/exambuddy/src/main/ExamBuddy.Api.Tests/Features/QuestionReviews/QuestionReviewOllamaLocalModel.cs)

Here is the pull request containing the complete set of changes. Note that it was implemented using the application’s previous architecture, before I decided to restructure it.

[Merge feature/ai integraion review questions and answers](https://bitbucket.org/mat-czernek/exambuddy/pull-requests/1)

If you would like to see the changes I made to align the application with Vertical Slice Architecture, check out this pull request.

[Merge feature/revisit-app-architecture](https://bitbucket.org/mat-czernek/exambuddy/pull-requests/2)

## Models tests

Selecting the Best Open-Source Model

*   gemma3:4b
    
*   phi4-mini:3.8b
    
*   aya-expanse:8b
    
*   ministral-3:8b
    
*   gemma3:12b-it-qat
    

Here is why I selected them:

*   gemma3:4b - this model is relatively small, at around 3.4 GB, so it can fit within my laptop’s 8 GB of VRAM. It also supports more than 140 languages.
    
*   gemma3:12b-it-qat - this model is larger than gemma3:4b, but the QAT variant reduces its VRAM requirements, making it more suitable for my modest laptop.
    
*   phi4-mini:3.8b - this model was trained on filtered, publicly available web data. It offers strong reasoning capabilities and performs well on mathematical and logical tasks.
    
*   ministral-3:8b - designed for local execution, this model supports system prompts and structured JSON responses. Based on its description alone, it was already a strong candidate.
    
*   aya-expanse:8b - one of its key advantages is its support for the Polish language.
    

It is also worth mentioning that the models were running on my gaming laptop, which was connected to the same local network as the machine from which I ran the evaluation. To learn how to configure Ollama on your local network, [see the following blog post.](https://mateusz-czernek.pl/make-ai-models-available-in-your-home-network-second-life-for-a-gaming-laptop)

The list was based on my own research into the available models, including their descriptions on the Ollama website and information I found elsewhere online. It is by no means representative, and someone with more experience working with LLMs might choose a completely different set of models. However, it was sufficient for my purposes—and for having some fun along the way.

The most suitable model may also vary depending on the question category. Since the application is currently used only by my wife and occasionally by my son, this approach works well enough for now. If necessary, I can also switch to a paid model relatively easily in the future.

For this small experiment, I prepared ten questions from different categories, along with their answers. In some cases, the question itself was incorrect; in others, correct answers were deliberately marked as incorrect.

The full list of questions is available [here](https://bitbucket.org/mat-czernek/exambuddy/src/main/ai-evaluation-reports/20261007-194725-065-question-review-v1/question-review-models-screening-test-v1.json). I then asked Codex to build an evaluation tool that uses the same question-review service as the ExamBuddy application. The tool is available [here](https://bitbucket.org/mat-czernek/exambuddy/src/main/ExamBuddy.AiEval/).

The tool runs the same prompt used by the ExamBuddy application against the provided questions and answers. Once the evaluation is complete, it generates an HTML report that allows the results to be reviewed and rated manually. I chose this approach because evaluating the output using only fixed assertions would be difficult, given the non-deterministic nature of LLM responses.

Based on the experiment results and my manual evaluation, I selected **ministral-3:8b** as the best general-purpose model for my needs.

The report and its manual ratings are available [here.](https://bitbucket.org/mat-czernek/exambuddy/src/main/ai-evaluation-reports/20261007-194725-065-question-review-v1/)

The following chart shows the models’ ratings based on my manual evaluation of their responses:

![](https://cdn.hashnode.com/uploads/covers/64d3deaccda572f1d04ef00a/eeb3b7ce-665f-477f-a93e-a93057ec7558.png align="center")

, here shows how long each model took to generate a response:

![](https://cdn.hashnode.com/uploads/covers/64d3deaccda572f1d04ef00a/86ec8af2-1f6d-4ed7-b24c-d4e1595e54f2.png align="center")

Although the response time was far from ideal, I preferred the suggestions generated by ministral-3:8b. They were detailed, accurate, and well structured. That was the trade-off I was willing to make: slower responses in exchange for higher-quality summaries and suggestions.

Either way, it was great fun to spend some time coding, experimenting, and figuring out which model best suited my needs.

I’m glad that I took the time to configure Ollama on my gaming laptop and make the models available across my home network. This allowed me to experiment with the integration at no additional cost. During the early stages of development, I ran the integration tests against Ollama many times; with a paid service, I could have burned through a considerable number of tokens very quickly.

This setup will not replace powerful paid models, but it is an excellent way to prototype an application, validate the integration, and gain confidence in the solution before investing money in it.
