# Keyed Masking for Secure Multi-User Semantic Communication Reproducibility package for the manuscript *Mask-as-Key Secure Multiple Access for Semantic Communications: Physical-Layer Encryption and Jamming Robustness*, submitted to the IEEE Transactions on Information Forensics and Security. Every figure and table in the paper is regenerated from this repository. Experiment scripts write CSV files to `data/` and never draw; `replot_security.py` reads only `data/` and writes the figure PDFs to `fig/`; `make_tables.py` prints the LaTeX rows of the result tables. ## Idea Multi-user semantic communication superposes several users on one frame and multiplies each user embedding by a distinct pattern so that the receivers can separate them. This code treats that pattern as a secret key. One keyed operation then encrypts each user against a receiver without the key, separates the users, and spreads a jammer that does not hold the key, at no extra bandwidth, power, or rate. ## Layout ``` code/ sse_lib.py transmit and receive core, channel, training, OMA reference exp_full.py stages A-F and L: SNR sweep, key length, jamming across schemes, key families, scheme comparison, attack difficulty exp_kpa.py stage H: known-plaintext attack on the key exp_refresh.py stage K: the key-refresh layer, invariance group exp_real_sec.py stage G: real BERT WordPiece token streams verify_math.py closed-form checks V1-V5 against Monte Carlo, PASS/FAIL replot_security.py every result figure, from data/ to fig/ make_tables.py LaTeX rows of every result table, from data/ feasibility_security.py early CPU-sized study, kept for the record data/ CSV results, one file per stage fig/ figure PDFs, regenerated by replot_security.py ``` ## Reproducing Requires Python 3, PyTorch, NumPy, and Matplotlib. The real-token stage additionally needs `datasets` and `transformers`. A CUDA device is recommended; the code falls back to CPU. Under Windows install PyTorch in WSL, because the native Windows build does not load the CUDA libraries. ```bash python verify_math.py # closed-form verification, prints PASS/FAIL python exp_full.py # stages A-F and L python exp_kpa.py # known-plaintext attack python exp_refresh.py # the key-refresh layer python exp_real_sec.py # real token streams python replot_security.py # all figures from the CSVs python make_tables.py # LaTeX rows of the result tables ``` Seeds are fixed: training 1, evaluation 777, attacker key guess 20260813, key recovery 4242, brute-force search 31, cross-scheme comparison 11, key refresh 5150. Re-running reproduces the released CSV files. ## Figure and table map | Artifact | Script | Data | |---|---|---| | Fig. 2 SER against SNR | `exp_full.stage_A` | `sec_snr.csv` | | Fig. 3 key length | `exp_full.stage_B` | `sec_keylen.csv` | | Fig. 4 jamming | `exp_full.stage_L` | `sec_jam_cmp.csv`, `sec_jam.csv` | | Fig. 5 key sensitivity | `exp_full.stage_I` | `sec_sens_cmp.csv` | | Fig. 6 brute-force search | `exp_full.stage_J` | `sec_brute_cmp.csv`, `sec_brute.csv` | | Fig. 7 known-plaintext attack | `exp_kpa` | `kpa.csv` | | Fig. 8 real token streams | `exp_real_sec` | `real_sec_ter.csv` | | Scheme comparison table | `exp_full.stage_E` | `sec_compare.csv` | | Key family table | `exp_full.stage_D` | `sec_maskfam.csv`, `sec_regjam.csv` | | Headline recovery table | `exp_real_sec` | `real_sec_stats.json` | | Key refresh tables | `exp_refresh` | `refresh_summary.csv`, `refresh_kpa.csv` | ## Security scope The analysis covers an adversary that observes transmitted frames. The masking is linear, so an adversary that also learns the indices some frames carried recovers the key from a few frames, which `exp_kpa.py` measures. The key must therefore be refreshed per coherence block from a shared seed. `exp_refresh.py` implements that layer and shows why it has to draw from the transformations that leave the decision statistic invariant: a refresh that installs fresh orthogonal keys instead costs the legitimate users a factor of nearly three, while the invariant refresh is free and raises the per-block key to 64.8 bits. ## License MIT for the code. The AG News data and the language-model tokenizer are obtained from their own sources under their own terms.