Swish-T: Enhancing Swish Activation With Tanh-Based Bias for Improved Neural Network Performance

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

We present the Swish-T family of activation functions, which extends Swish by integrating a bounded, zero-centered Tanh-based bias term inside the activation. This design provides finer control near the activation threshold while preserving computational simplicity as a drop-in replacement. We evaluate Swish-T on diverse benchmarks, including MNIST, Fashion-MNIST, SVHN, CIFAR-10/100, Tiny-ImageNet, and Cityscapes, covering image classification and semantic segmentation across 12 architectures (CNNs and a transformer baseline). Across these settings, Swish-T consistently matches or improves upon widely used activations such as ReLU, GELU, and Swish, while offering a more efficient alternative to SMU. For example, replacing ReLU with Swish-TC in ShuffleNetV2 on CIFAR-100 improves Top-1 accuracy by 4.12%, and replacing ReLU with Swish-T in PRN-50 on Tiny-ImageNet improves accuracy by 0.97%. Compared to SMU, which can incur substantial training-time and memory overhead, Swish-T achieves comparable or better accuracy with lower computational cost, making it a practical activation choice for a broad range of deep learning models.

키워드

Computational efficiencyOptimizationVisualizationTrainingStability analysisShapeImage classificationDeep learningComputer architectureComputational modelingActivation functiondeep learningimage classificationsemantic segmentationmachine learningcomputer visionoptimization
제목
Swish-T: Enhancing Swish Activation With Tanh-Based Bias for Improved Neural Network Performance
저자
Seo, YoungminKim, JinhaPark, Unsang
DOI
10.1109/ACCESS.2026.3667968
발행일
2026-02
유형
Article
저널명
IEEE Access
14
페이지
34404 ~ 34419