Unconstrained key training converged to disjoint sparse supports: 99 percent of each users key energy sat on three or four of the sixteen entries, with pairwise disjoint supports and one numerically dead codebook column. That is an orthogonal slot allocation, so the superposition collapsed into OMA and the key space was far smaller than the dense direction the brute-force study assumes. The main configuration is now the structured Walsh-Hadamard family, which is dense, exactly orthogonal, unit modulus, and already the best family in the key-family table. base_keys generalizes to any key length by truncating the next power-of-two Sylvester order, and the key-length sweep keeps only lengths where the truncated rows stay exactly orthogonal, verified numerically. Also fixes the M-PAM energy normalization in oma_ser_keylen, which used sqrt(6g/(M^2-1)) where unit average symbol energy gives A^2=3/(M^2-1); the closed form was 3 dB optimistic and now reproduces a direct Monte Carlo to 1e-5. Results move accordingly: the proposal now stays below OMA at every SNR and reaches 1.52x at key length 64, while the jamming margin falls to 5.5-6.3 dB and the brute-force curve to 0.59 at a million guesses.
Keyed Masking for Secure Multi-User Semantic Communication
Reproducibility package for the manuscript Mask-as-Key Secure Multiple Access for Semantic Communications: Physical-Layer Encryption and Jamming Robustness, submitted to the IEEE Transactions on Information Forensics and Security.
Every figure and table in the paper is regenerated from this
repository. Experiment scripts write CSV files to data/ and never
draw; replot_security.py reads only data/ and writes the figure PDFs
to fig/; make_tables.py prints the LaTeX rows of the result tables.
Idea
Multi-user semantic communication superposes several users on one frame and multiplies each user embedding by a distinct pattern so that the receivers can separate them. This code treats that pattern as a secret key. One keyed operation then encrypts each user against a receiver without the key, separates the users, and spreads a jammer that does not hold the key, at no extra bandwidth, power, or rate.
Layout
code/
sse_lib.py transmit and receive core, channel, training, OMA reference
exp_full.py stages A-F and L: SNR sweep, key length, jamming across
schemes, key families, scheme comparison, attack difficulty
exp_kpa.py stage H: known-plaintext attack on the key
exp_refresh.py stage K: the key-refresh layer, invariance group
exp_permkpa.py permutation-key known-plaintext attack (Fig. 7)
check_cov_*.py ciphertext-only covariance-attack checks (referee M1)
exp_real_sec.py stage G: real BERT WordPiece token streams
verify_math.py closed-form checks V1-V5, PASS/FAIL and verify_math.csv
replot_security.py every result figure, from data/ to fig/
make_tables.py LaTeX rows of every result table, from data/
feasibility_security.py early CPU-sized study, kept for the record
data/ CSV results, one file per stage
fig/ figure PDFs, regenerated by replot_security.py
Reproducing
Requires Python 3, PyTorch, NumPy, and Matplotlib. The real-token stage
additionally needs datasets and transformers. A CUDA device is
recommended; the code falls back to CPU. Under Windows install PyTorch
in WSL, because the native Windows build does not load the CUDA
libraries.
python verify_math.py # closed-form verification, prints PASS/FAIL
python exp_full.py # stages A-F and L
python exp_kpa.py # known-plaintext attack
python exp_refresh.py # the key-refresh layer
python exp_real_sec.py # real token streams
python replot_security.py # all figures from the CSVs
python make_tables.py # LaTeX rows of the result tables
Seeds are fixed: training 1, evaluation 777, attacker key guess 20260813, key recovery 4242, brute-force search 31, cross-scheme comparison 11, key refresh 5150. Re-running reproduces the released CSV files.
Figure and table map
| Artifact | Script | Data |
|---|---|---|
| Fig. 2 SER against SNR | exp_full.stage_A |
sec_snr.csv |
| Fig. 3 key length | exp_full.stage_B |
sec_keylen.csv |
| Fig. 4 jamming (4 schemes) | exp_full.stage_L |
sec_jam_cmp.csv, sec_jam.csv |
| Fig. 5 key sensitivity | exp_full.stage_I |
sec_sens_cmp.csv |
| Fig. 6 brute-force search | exp_full.stage_J |
sec_brute_cmp.csv, sec_brute.csv |
| Fig. 7 known-plaintext attack | exp_kpa, exp_permkpa |
kpa.csv, pkpa.csv |
| Fig. 8 real token streams | exp_real_sec |
real_sec_ter.csv |
| Scheme comparison table | exp_full.stage_E |
sec_compare.csv |
| Key family table | exp_full.stage_D |
sec_maskfam.csv, sec_regjam.csv |
| Headline recovery table | exp_real_sec |
real_sec_stats.json |
| Key refresh tables | exp_refresh |
refresh_summary.csv, refresh_kpa.csv |
Security scope
The analysis covers an adversary that observes transmitted frames. The
masking is linear, so an adversary that also learns the indices some
frames carried recovers the key from a few frames, which exp_kpa.py
measures. The key must therefore be refreshed per coherence block from a shared
seed. exp_refresh.py implements that layer and shows why it has to
draw from the transformations that leave the decision statistic
invariant: a refresh that installs fresh orthogonal keys instead costs
the legitimate users a factor of nearly three, while the invariant
refresh costs nothing and raises the per-block key from 15.0 to 64.8
bits.
License
MIT for the code. The AG News data and the language-model tokenizer are obtained from their own sources under their own terms.