This study aims to examine the capability of Large Language Models (LLMs) to generate stock investment recommendations in the form of buy, hold, and sell for stocks included in the LQ45 Index of the Indonesia Stock Exchange. A quantitative experimental approach was employed using three generative artificial intelligence models: ChatGPT, Gemini, and DeepSeek. Each model received an identical prompt incorporating stock price information, trading volume, candlestick patterns, and macroeconomic and microeconomic sentiment. The research sample comprised 45 LQ45 constituent stocks observed over 20 trading days in December 2025, resulting in 900 observations for each model and 2,700 observations in total. Model performance was evaluated based on three primary dimensions: recommendation accuracy relative to actual stock price movements, investment returns, and downside risk. The findings indicate discrepancies between LLM-generated recommendations and actual stock price movements, suggesting that none of the three models can predict stock price direction with complete accuracy. Nevertheless, all models generated positive cumulative returns. DeepSeek achieved the highest Compounded Cumulative Return (CCR) at 24.01%, followed by Gemini at 23.58% and ChatGPT at 1.90%. In terms of risk, ChatGPT demonstrated the most favorable performance, recording the lowest total risk at 16.35%, compared with DeepSeek at 22.34% and Gemini at 24.85%. The loss rates were also below 50% across all models, namely 34.11% for DeepSeek, 37.88% for Gemini, and 27.33% for ChatGPT. These findings suggest that LLMs can function as Decision Support Systems for investment decision-making by generating recommendations with the potential to produce positive returns while maintaining relatively moderate risk. However, LLM-generated recommendations should not be interpreted as definitive market forecasts because their performance remains sensitive to market volatility, information dynamics, and rapidly changing financial conditions.
Copyrights © 2026