ResearchDemo offline · 2024 – 2026

Image Captioning System — CNN + Transformer

The code behind a co-authored IEEE conference paper on image captioning, lifted from a research notebook into a typed, tested, configuration-driven Python package with a FastAPI inference service and a React 19 front end. The trained model is versioned on the Hugging Face Hub, and a pre-registered evaluation audit tests whether its scores mean what they appear to.

Measured, not claimed

10.4 BLEU-4

Beam search, as committed

500 COCO images · ~1.5 references each

25.9 BLEU-4

Same predictions, five-reference rescore

sacreBLEU · identical 500 predictions

3 of 30

Captions judged image-specific

Blinded rubric review, judged without sight of the BLEU result

94 tests

Four-job CI

ruff · mypy · notebook freeze · frontend build

Image Captioning System — CNN + Transformer preview

Technology stack

PythonTensorFlow / KerasInceptionV3TransformerFastAPIPydantic v2React 19pytestmypyHugging Face Hub
01

The paper

“AI Narratives: Bridging Visual Content and Linguistic Expression”, 2024 IEEE International Conference on Smart Power Control and Renewable Energy (ICSPCRE). Authors: Preetam, Sai Chetan Muppalla, Apoorv Raj and Jasneet Chawla. It applies an InceptionV3 image encoder and a Transformer decoder to COCO image captioning.

02

From notebook to package

  • YAML configuration validated by Pydantic v2 (unknown keys are rejected) and overridable through environment variables. Hyperparameters mirror the notebook: COCO 2017, 120,000 sampled captions, a 15,000-token vocabulary, 40-token captions and an 80/20 split.
  • Frozen ImageNet InceptionV3 features feed a Transformer encoder and a single Transformer decoder layer with 8 attention heads.
  • The original research notebook is frozen by SHA-256 and checked in CI, so the reference implementation cannot drift silently.
  • 94 automated tests. CI runs ruff and mypy, a Python test matrix, the notebook freeze check and a frontend build.
03

Serving

  • A FastAPI service with a lifespan-managed predictor, so one warm model is reused across requests, plus request-scoped structured logging and a multipart POST /v1/captions endpoint.
  • A React 19 + Vite + Tailwind front end with request timeouts and structured error handling.
  • Trained weights (v2.0.0) versioned on the Hugging Face Hub.
04

The evaluation audit

A reproducible harness computes BLEU-1–4, CIDEr, METEOR and ROUGE-L for greedy and beam-search decoding and writes per-run metrics, predictions and diagnostics, so two runs can be compared mechanically.

Scored against about 1.5 reference captions per image, the v2.0.0 model reaches BLEU-4 10.4 with beam search (10.6 greedy). A pre-registered rescore of the same 500 predictions against all five COCO references gives 25.9 (sacreBLEU; NLTK variants give 22.2 and 23.8). The low single-reference score was mostly a property of the scoring setup, not the model.

A blinded review of 30 captions, run separately so the BLEU number could not bias it, found 3 image-specific and correct, 11 generic, 15 partially correct and 1 wrong. A good corpus score does not mean the captions are good, so the audit’s decision rule flags the result for human review instead of declaring success.

05

Status

Paper published in 2024; the package, service and evaluation work followed in 2026. The hosted inference Space is currently offline because of a deployment configuration error since June 2026, so there is no live demo. Baselines against other captioning models, latency benchmarks and observability are on the roadmap, not done.

Sources

Checked against the repositories on 22 September 2026.