Image Captioning System — CNN + Transformer
The code behind a co-authored IEEE conference paper on image captioning, lifted from a research notebook into a typed, tested, configuration-driven Python package with a FastAPI inference service and a React 19 front end. The trained model is versioned on the Hugging Face Hub, and a pre-registered evaluation audit tests whether its scores mean what they appear to.
Measured, not claimed
10.4 BLEU-4
Beam search, as committed
500 COCO images · ~1.5 references each
25.9 BLEU-4
Same predictions, five-reference rescore
sacreBLEU · identical 500 predictions
3 of 30
Captions judged image-specific
Blinded rubric review, judged without sight of the BLEU result
94 tests
Four-job CI
ruff · mypy · notebook freeze · frontend build
Technology stack
The paper
“AI Narratives: Bridging Visual Content and Linguistic Expression”, 2024 IEEE International Conference on Smart Power Control and Renewable Energy (ICSPCRE). Authors: Preetam, Sai Chetan Muppalla, Apoorv Raj and Jasneet Chawla. It applies an InceptionV3 image encoder and a Transformer decoder to COCO image captioning.
From notebook to package
- YAML configuration validated by Pydantic v2 (unknown keys are rejected) and overridable through environment variables. Hyperparameters mirror the notebook: COCO 2017, 120,000 sampled captions, a 15,000-token vocabulary, 40-token captions and an 80/20 split.
- Frozen ImageNet InceptionV3 features feed a Transformer encoder and a single Transformer decoder layer with 8 attention heads.
- The original research notebook is frozen by SHA-256 and checked in CI, so the reference implementation cannot drift silently.
- 94 automated tests. CI runs ruff and mypy, a Python test matrix, the notebook freeze check and a frontend build.
Serving
- A FastAPI service with a lifespan-managed predictor, so one warm model is reused across requests, plus request-scoped structured logging and a multipart POST /v1/captions endpoint.
- A React 19 + Vite + Tailwind front end with request timeouts and structured error handling.
- Trained weights (v2.0.0) versioned on the Hugging Face Hub.
The evaluation audit
A reproducible harness computes BLEU-1–4, CIDEr, METEOR and ROUGE-L for greedy and beam-search decoding and writes per-run metrics, predictions and diagnostics, so two runs can be compared mechanically.
Scored against about 1.5 reference captions per image, the v2.0.0 model reaches BLEU-4 10.4 with beam search (10.6 greedy). A pre-registered rescore of the same 500 predictions against all five COCO references gives 25.9 (sacreBLEU; NLTK variants give 22.2 and 23.8). The low single-reference score was mostly a property of the scoring setup, not the model.
A blinded review of 30 captions, run separately so the BLEU number could not bias it, found 3 image-specific and correct, 11 generic, 15 partially correct and 1 wrong. A good corpus score does not mean the captions are good, so the audit’s decision rule flags the result for human review instead of declaring success.
Status
Paper published in 2024; the package, service and evaluation work followed in 2026. The hosted inference Space is currently offline because of a deployment configuration error since June 2026, so there is no live demo. Baselines against other captioning models, latency benchmarks and observability are on the roadmap, not done.
Sources
- Source code
- Paper on IEEE Xplore
- DOI 10.1109/ICSPCRE62303.2024.10675203
- Model weights on the Hugging Face Hub
- Evaluation audit verdict
Checked against the repositories on 22 September 2026.