Files
TOIFAS/README.md
T
KiHoLee 3529ab1918 Audit round: fair OMA reference, dense grids, covariance-attack checks
Resource-match the OMA reference in the key-length sweep (oma_ser_keylen),
which gives it the L/16 combining gain the longer frame allows. The
proposal now passes a resource-matched OMA by 1.27x at L=64 rather than
the 4.3x reported against a fixed-d reference.

Densify the JSR, sensitivity, and brute-force grids so the curves are
smooth, give the index cipher its channel floor instead of error-free
reception, and add the permutation-key known-plaintext attack
(exp_permkpa) so Fig. 7 carries a conventional linear scheme.

Add check_cov_attack.py and check_cov_ceiling.py: a referee raised a
ciphertext-only second-order attack; the exact-population test shows the
received covariance leaks only a sparse rank-deficient subset of the key
Gram and leaves the eavesdropper at the random-guess level.

Dump verify_math.csv, move the superseded V=256 pilot CSVs to data/pilot.
2026-08-16 23:48:39 +09:00

98 lines
4.5 KiB
Markdown

# Keyed Masking for Secure Multi-User Semantic Communication
Reproducibility package for the manuscript *Mask-as-Key Secure Multiple
Access for Semantic Communications: Physical-Layer Encryption and
Jamming Robustness*, submitted to the IEEE Transactions on Information
Forensics and Security.
Every figure and table in the paper is regenerated from this
repository. Experiment scripts write CSV files to `data/` and never
draw; `replot_security.py` reads only `data/` and writes the figure PDFs
to `fig/`; `make_tables.py` prints the LaTeX rows of the result tables.
## Idea
Multi-user semantic communication superposes several users on one frame
and multiplies each user embedding by a distinct pattern so that the
receivers can separate them. This code treats that pattern as a secret
key. One keyed operation then encrypts each user against a receiver
without the key, separates the users, and spreads a jammer that does not
hold the key, at no extra bandwidth, power, or rate.
## Layout
```
code/
sse_lib.py transmit and receive core, channel, training, OMA reference
exp_full.py stages A-F and L: SNR sweep, key length, jamming across
schemes, key families, scheme comparison, attack difficulty
exp_kpa.py stage H: known-plaintext attack on the key
exp_refresh.py stage K: the key-refresh layer, invariance group
exp_permkpa.py permutation-key known-plaintext attack (Fig. 7)
check_cov_*.py ciphertext-only covariance-attack checks (referee M1)
exp_real_sec.py stage G: real BERT WordPiece token streams
verify_math.py closed-form checks V1-V5, PASS/FAIL and verify_math.csv
replot_security.py every result figure, from data/ to fig/
make_tables.py LaTeX rows of every result table, from data/
feasibility_security.py early CPU-sized study, kept for the record
data/ CSV results, one file per stage
fig/ figure PDFs, regenerated by replot_security.py
```
## Reproducing
Requires Python 3, PyTorch, NumPy, and Matplotlib. The real-token stage
additionally needs `datasets` and `transformers`. A CUDA device is
recommended; the code falls back to CPU. Under Windows install PyTorch
in WSL, because the native Windows build does not load the CUDA
libraries.
```bash
python verify_math.py # closed-form verification, prints PASS/FAIL
python exp_full.py # stages A-F and L
python exp_kpa.py # known-plaintext attack
python exp_refresh.py # the key-refresh layer
python exp_real_sec.py # real token streams
python replot_security.py # all figures from the CSVs
python make_tables.py # LaTeX rows of the result tables
```
Seeds are fixed: training 1, evaluation 777, attacker key guess
20260813, key recovery 4242, brute-force search 31, cross-scheme
comparison 11, key refresh 5150. Re-running reproduces the released CSV
files.
## Figure and table map
| Artifact | Script | Data |
|---|---|---|
| Fig. 2 SER against SNR | `exp_full.stage_A` | `sec_snr.csv` |
| Fig. 3 key length | `exp_full.stage_B` | `sec_keylen.csv` |
| Fig. 4 jamming (4 schemes) | `exp_full.stage_L` | `sec_jam_cmp.csv`, `sec_jam.csv` |
| Fig. 5 key sensitivity | `exp_full.stage_I` | `sec_sens_cmp.csv` |
| Fig. 6 brute-force search | `exp_full.stage_J` | `sec_brute_cmp.csv`, `sec_brute.csv` |
| Fig. 7 known-plaintext attack | `exp_kpa`, `exp_permkpa` | `kpa.csv`, `pkpa.csv` |
| Fig. 8 real token streams | `exp_real_sec` | `real_sec_ter.csv` |
| Scheme comparison table | `exp_full.stage_E` | `sec_compare.csv` |
| Key family table | `exp_full.stage_D` | `sec_maskfam.csv`, `sec_regjam.csv` |
| Headline recovery table | `exp_real_sec` | `real_sec_stats.json` |
| Key refresh tables | `exp_refresh` | `refresh_summary.csv`, `refresh_kpa.csv` |
## Security scope
The analysis covers an adversary that observes transmitted frames. The
masking is linear, so an adversary that also learns the indices some
frames carried recovers the key from a few frames, which `exp_kpa.py`
measures. The key must therefore be refreshed per coherence block from a shared
seed. `exp_refresh.py` implements that layer and shows why it has to
draw from the transformations that leave the decision statistic
invariant: a refresh that installs fresh orthogonal keys instead costs
the legitimate users a factor of nearly three, while the invariant
refresh costs nothing and raises the per-block key from 15.0 to 64.8
bits.
## License
MIT for the code. The AG News data and the language-model tokenizer are
obtained from their own sources under their own terms.