<?xml version="1.0" encoding="UTF-8"?>
<article article-type="research-article" xml:lang="en" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher">global-journal-of-computer-science-and-technology-d-neural-ai</journal-id>
<journal-title-group>
<journal-title>Global Journal of Computer Science and Technology - D: Neural &amp; AI</journal-title>
</journal-title-group>
<issn publication-format="print">0975-4350</issn>
<issn publication-format="electronic">0975-4172</issn>
<publisher><publisher-name>Global Journals Publishing Group Incorporated</publisher-name></publisher>
<self-uri xlink:href="https://globaljournals.org/journal-seo-export/jats/259174.xml" />
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.34257/GJCSTD259174</article-id>
<article-id pub-id-type="publisher-id">259174</article-id>
<title-group>
<article-title>Vision Language Models as the Reasoning Layer of an AI-CCTV Planning Toolkit: A Case Study of the IndoAI’s Calculator Suite</article-title>
<subtitle>VLM as AI-CCTV Planning Reasoning Layer</subtitle>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Gujar</surname><given-names>Vivek</given-names></name><xref ref-type="aff" rid="aff1" />
</contrib>
<contrib contrib-type="author"><name><surname>Gujar</surname><given-names>Advait</given-names></name><xref ref-type="aff" rid="aff2" />
</contrib>
<contrib contrib-type="author"><name><surname>Rathore</surname><given-names>Ashwani</given-names></name></contrib>
</contrib-group>
<aff id="aff1">INDIA, IndoAI Technologies Pvt. Ltd.</aff>
<aff id="aff2">INDIA, Manipal Institute of Technology</aff>
<pub-date publication-format="electronic" date-type="pub" iso-8601-date="2026-08-01">
<day>01</day>
<month>08</month>
<year>2026</year>
</pub-date>
<volume>26</volume>
<abstract><p>Designing AI-enabled video surveillance systems has become far more challenging than simply deciding the number of cameras or estimating storage. Once AI inference is performed on cameras or edge devices, every design decision influences several others. For example, increasing the pixel density needed for facial recognition affects lens selection, bitrate, storage capacity, and network bandwidth, while adding more AI analytics changes edge-computing requirements and software licensing. As a result, conventional CCTV planning methods are no longer sufficient for modern AI deployments. This paper presents the architecture of IndoAI’s integrated planning toolkit, which follows a three-stage workflow ie Design, Size, and Verify, using nine specialised calculators for camera placement, infrastructure sizing, and deployment validation. Alongside these tools, an Appization-based AI Agent License Sizing and ROI Predictor estimates software licensing requirements and deployment economics. Building on recent advances in Vision Language Models (VLMs) and platform economics, we propose a VLM as the natural-language interface to this planning framework. Users can describe surveillance requirements using text, photographs or floor plans instead of manually entering technical parameters. The VLM converts these inputs into structured data for the calculator suite. We also present two systems under development: a VLM-driven video intelligence pipeline that analyses surveillance footage to generate timestamped event detections, evaluated using the UCF-Crime dataset and a VLM-assisted Bill of Quantities generator that produces annotated camera layouts together with storage, bandwidth and licensing estimates for deployment planning.</p></abstract>
<kwd-group kwd-group-type="author-generated">
<kwd>Vision Language Models</kwd>
<kwd>AI video surveillance</kwd>
<kwd>Bill of Quantities automation</kwd>
<kwd>edge computing</kwd>
<kwd>platform economics</kwd>
<kwd>pixel density planning</kwd>
<kwd>license/agent provisioning</kwd>
<kwd>Indoai.</kwd>
</kwd-group>
<self-uri content-type="pdf" xlink:href="https://globaljournals.org:/GJCST_Volume26/vision-language-models-as-the-reasoning-layer-of-an-ai-cctv-p-4d39273165.pdf?v=1784203302498#" />
<self-uri content-type="html" xlink:href="https://globaljournals.org/scholarly-articles/vlm-as-ai-cctv-planning-reasoning-layer/" />
</article-meta>
</front>
<body>
<sec>
<title>Full Text</title>
<p></p>
</sec>
</body>
</article>