Publication Date

Spring 2026

Degree Type

Master's Project

Degree Name

Master of Science in Computer Science (MSCS)

Department

Computer Science

First Advisor

Teng Moh

Second Advisor

Melody Moh

Third Advisor

Maryam Khazaei

Keywords

Large Language Model, Natural Language Processing, Benchmark, Chain-of-Thought

Abstract

Hallucinations, plausible yet incorrect, are prevalent across large language models undermining confidence in their reliability. This study investigates mitigation approaches for hallucinations in large language models. The study examines its effectiveness by using code generation tasks as a benchmark. Using 141 coding problems, the study compares zero-shot inference, Chain-of-Thought, and a hybrid Chain-of-Thought approach that incorporates review, optimization, and testing phases. Four large language models that were evaluated through the different approaches were Llama 3.3 and Gemma 2 (general-purpose models), DeepSeek R1 (internal Chain-of-Thought), and Qwen 2.5 Coder (fine-tuned model). Evaluations take into account accuracy, token utilization, and generation time. The results demonstrate that reasoning-enhanced approaches consistently improve accuracy around 3% to 15%, with the hybrid Chain-of-Thought methodology showing the most significant gains. The findings suggest that targeted prompting strategies encouraging reasoning, review, and testing can significantly enhance the reliability of LLM-generated code, with consistent improvements observable across different models.

Available for download on Saturday, May 22, 2027

Share

COinS