From my own experience, when using the Gemma 4 31B model with the same network latency (35-45ms), I was able to get a faster and more accurate response. However, the Gemini 3.1 Flash Lite model would make me wait for 4-5 seconds on average before receiving the answer, and the result felt somewhat unsatisfactory.
I don't know why the Kagi officials decided to switch to this model. Personally, I think it's to reduce costs and improve efficiency; but it is obvious that the latter's answer is not as good as the former's.
Let's put aside the issue of the decline in answer quality for now, and I hope that the Kagi team will make some adaptations to the LLM of KQA.
BTW & FYI, I usually call KQA = Kagi (AI) Quick Answer, which is convenient.
I think it's better to allow users to choose which LLM of KQA they want to use.
(I think this feature is both belonging to "Kagi Search" and "Kagi Assistant")