Abstract
The exponential growth of academic publishing has intensified the methodological demands placed on systematic literature reviews (SLRs), creating a pressing tension between comprehensiveness and feasibility. While SLRs remain a cornerstone of knowledge synthesis in management and business research, their execution has become increasingly burdensome due to expanding publication volumes, terminological heterogeneity, and the cognitive load associated with manual screening and classification. Concurrently, Large Language Models (LLMs) and generative AI tools are reshaping academic workflows, yet their integration into SLR methodology lacks structured, transparent, and reproducible guidance. Existing discourse largely highlights potential benefits without specifying operational procedures, validation mechanisms, or reporting standards.
This developmental paper addresses that gap by proposing a structured hybrid framework that embeds LLM-based tools and rule-based automation within an established systematic review logic. The framework is organized across three phases. The first phase covers planning and search design, including domain familiarization, research gap verification, keyword development, Boolean query construction, and protocol development, where LLMs support language-intensive tasks such as synonym generation, query refinement, and structured drafting while human oversight governs all conceptual decisions. The second phase addresses execution and screening, where deterministic tasks such as deduplication and journal quality assessment are assigned to rule-based automation, and LLM-assisted classification supports title, abstract, and keyword screening under explicitly designed prompts with mandatory validation procedures. The third phase encompasses data extraction, synthesis, and reporting, where AI tools accelerate structured capture and assist thematic clustering, while researcher judgment remains central to interpretive accuracy and transparent documentation.
A central methodological contribution of the framework is its principled distinction between tasks suited to generative AI, those best handled through deterministic automation, and those that must remain exclusively human-led. Full-text quality assessment and final inclusion decisions, for instance, are designated as manual activities given their interpretive demands, while overlap validation and journal ranking matching benefit from scripted automation due to their logical determinism. This alignment of tool selection with task characteristics ensures that efficiency gains do not compromise methodological rigor.
By situating AI as an augmentation mechanism rather than a decision-making agent, the framework advances toward more scalable evidence synthesis without sacrificing transparency or replicability. Prompts, validation logic, and executable code accompany the procedural design, enabling scholarly audit and replication. The paper contributes to the emerging methodological discourse on responsible AI integration in academic research and invites critical engagement, refinement, and empirical calibration from the research community.