Kai Zhang
Financial Engineering, Baruch College, NY, USA

Published : 2 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 2 Documents
Search

News-Based Uncertainty and Macro-Market Fusion for VIX Direction Forecasting: Evidence from 2015-2024 FRED Panel Hailin Zhou; Kai Zhang
Journal of Technology Informatics and Engineering Vol. 4 No. 2 (2025): AUGUST | JTIE : Journal of Technology Informatics and Engineering
Publisher : University of Science and Computer Technology

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.51903/jtie.v4i2.540

Abstract

This paper evaluates news-based uncertainty and macro-market fusion for one-trading-day VIX direction forecasting using a 2015-2024 daily FRED panel. The dependent variable equals one when the next trading-day VIX close exceeds the current close. The final panel contains 2,512 processed observations from 2015-02-04 to 2024-12-30, with 1,997 training observations from 2015-2022, 257 validation observations in 2023, and 258 holdout observations in 2024. The feature set combines VIX market-state variables, the daily newspaper-based Economic Policy Uncertainty index, the effective federal funds rate, 10-year and 2-year Treasury yields, lagged CPI inflation, lagged unemployment, and interaction terms. Expanding cross-validation on the 2015-2022 training sample gives the highest average ROC AUC to Fusion Random Forest (0.5620). The 2023 validation window selects a weighted fusion ensemble with ROC AUC 0.6005; its weights are fixed before the 2024 test. In the 2024 holdout window, VIX-only Logistic achieves the highest one-day ROC AUC (0.5869) and F1 (0.5448), while the weighted fusion ensemble reaches ROC AUC 0.5638 and F1 0.4786. Event-window diagnostics show that EPU shocks following calm VIX states have a next-day VIX-up rate of 0.5672, compared with 0.4445 on other trading days. The findings support a cautious interpretation: news-based uncertainty contains conditional information, but one-day practical forecasting reliability remains modest and VIX state variables remain the strongest 2024 holdout benchmark.
Numerical-Reasoning Guardrails for a Quant Research Assistant: A Compact Reproducible Benchmark Using SEC and FRED Data Zeyi Li; Kai Zhang; Annie Wong
Journal of Technology Informatics and Engineering Vol. 5 No. 2 (2026): AUGUST | JTIE : Journal of Technology Informatics and Engineering
Publisher : University of Science and Computer Technology

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.51903/jtie.v5i2.541

Abstract

This paper presents a compact reproducible benchmark for evaluating numerical-reasoning guardrails in a quant research assistant. The revised experiment uses a fixed 2026 source snapshot derived from the SEC 2026 Q1 Financial Statement Data Sets and FRED CSV series for VIXCLS, DGS10, DGS3MO and T10Y3M. The benchmark contains 460 tasks: 300 SEC financial-ratio tasks over 50 issuer-period records, 120 FRED VIX and Treasury-rate change tasks, and 40 macro-regime classification tasks. Each answer is evaluated by five programmatic guardrails: numeric consistency, unit correctness, time-window correctness, formula correctness and citation/source consistency. Four controlled response profiles are tested: Naive-RAG, Calculator-Only, Prompted-Checklist and Guarded-Quant. These profiles are deterministic failure-mode controls rather than performance claims about any particular deployed LLM. The empirical results show that arithmetic alone is not sufficient for financial safety: Calculator-Only reaches 79.78% numeric accuracy but only 0.43% all-guardrails pass rate because source, unit, formula and window fields often fail. Guarded-Quant achieves an 88.48% all-guardrails pass rate, 97.17% numeric accuracy, 100.00% unit pass rate, 96.30% window pass rate, 98.26% formula pass rate and 96.30% citation pass rate. The findings support a modest claim: a compact benchmark can make numerical audit failures visible, but it should not be read as evidence of broad quant-assistant reliability without broader data, live model outputs and operational stress tests.