Artificial intelligence, particularly generative AI, is increasingly positioned as an active participant in learning rather than a passive digital tool. This systematic literature review synthesizes evidence on human-AI collaborative problem solving in mathematics, with particular attention to AI roles, collaboration patterns, learning outcomes, and pedagogical orchestration. The review followed PRISMA 2020. A Scopus title-abstract-keyword search produced 369 source records; after 10 duplicates were removed, 359 records were screened, 125 full-text reports were assessed, and 10 studies were retained for qualitative synthesis. Because the original screening archive did not preserve reviewer-level decision logs, inter-rater agreement coefficients could not be retrospectively computed without reconstruction; this limitation and a prescribed dual-reviewer rerun protocol are reported explicitly. The synthesis identified four recurring AI roles: adaptive scaffold/tutor, cognitive or visualization tool, co-creative partner, and orchestration platform. Productive collaboration was most consistently associated with socially structured routines, shared agency, reflective activity, and explicit teacher orchestration, whereas unstructured use increased the risk of cognitive off-loading and over-delegation. Learning gains were reported for problem solving, conceptual understanding, motivation, self-efficacy, and related outcomes, but causal evidence remains limited by heterogeneous designs, levels, and measures. The review therefore treats orchestration—not model capability alone—as the principal design mechanism and recommends preregistered, longitudinal, middle-school studies with process-level measures, independent dual screening, and multi-database searching.