Export partially-captured groups as stable

Stop treating partially-captured oldest timestamp groups (including Arc 10-pulls) as a special "incomplete" case. All decoded rows are now exported and marked with stable UIDs; the previous INCOMPLETE_* warnings and related console hint logic were removed. annotate_groups and annotate_arc_groups were simplified to always set uid_status="stable" and export_record=True, and arc_stability_warnings was deleted. CLI, live capture runner, console, tests, docs and mappings were updated to reflect the new behavior and to remove emission/handling of those informational warnings.
This commit is contained in:
Golumpa 2026-06-11 14:57:08 +01:00
parent 20b2086299
commit 5d16dac4d5
10 changed files with 53 additions and 170 deletions

View file

@ -57,9 +57,9 @@ NTE history records do not appear to contain a unique server-side roll ID. UIDs
History always loads page 1 first and is scrolled downward, so the exporter anchors to the continuous run of pages starting at page 1 and ignores anything after the first gap (with a warning). This keeps the newest pages even if a later page is lost, and guarantees the newest timestamp group's ordinal 0 is captured.
Within a timestamp group, ordinal 0 is the newest record and unseen rows can only append after the captured ones, so every exported UID is stable. The one nuance is the oldest captured group: if the capture did not reach the true end of history and its dice-roll count is not a complete multiple of 10, it may be an unfinished 10-pull continuing onto an uncaptured page. Its captured prefix is still ordinal-stable, so it is exported with an informational warning telling you to scroll further to capture the rest.
Within a timestamp group, ordinal 0 is the newest record and unseen rows can only append after the captured ones, so every exported UID is stable -- including a partially captured oldest 10-pull. All decoded rows are therefore exported. Re-scanning later simply adds any rows that were not yet captured, with the same UIDs for the rows already seen.
For Monopoly, Points Gift and Chase Reward rows stay in the timestamp group for UID ordinal generation, but only `result_type = dice` rows count toward pull-set sizing. Arc pulls are always 10-pulls, so the same rule applies with a fixed group size of 10. In both systems the oldest captured group is exported even if it is an unfinished pull set (its captured prefix is ordinal-stable), and flagged with a warning so you know to scroll further on a later scan.
For Monopoly, Points Gift and Chase Reward rows stay in the timestamp group for UID ordinal generation, but only `result_type = dice` rows count toward pull-set sizing. Arc pulls are always 10-pulls. In both systems every captured group is exported, including the oldest one even if it is a partially captured pull set, because its captured prefix is ordinal-stable.
## Current Adapters

View file

@ -7,7 +7,7 @@ Known limitations:
- Other NTE banners are not implemented yet.
- The game appears not to provide a unique server-side roll ID in the decoded record body.
- UIDs are generated deterministically from decoded fields and timestamp-group order.
- An incomplete oldest timestamp group is still exported (its captured prefix has stable UIDs) and flagged so the user can scroll further on a later scan.
- A partially captured oldest timestamp group is still exported; its captured prefix has stable UIDs, and a later deeper scan adds the rest with the same UIDs.
- Pages are anchored to the continuous run starting at page 1; pages after the first gap are ignored and reported as warnings.
- Live capture currently uses Windows raw sockets and requires administrator permission.
- The file adapter reads mitmproxy `.flows` captures for research and testing.

View file

@ -36,7 +36,7 @@ Page and row numbers are research metadata only. They must not be used for perma
Timestamp groups keep all records with the same raw timestamp together for UID ordinal generation. For boundary/group-size detection, only `result_type = dice` rows count as pull-set members; Points Gift and Chase Reward rows stay in the group but do not increase the dice-only group count.
Pages are anchored to the continuous run starting at page 1 (history always loads page 1 first), so the newest timestamp group's ordinal 0 is always captured. Ordinals are assigned in scan order (newest first), so ordinal 0 of a timestamp group is its newest record and any unseen continuation rows can only append after the captured ones with higher ordinals. Every exported UID is therefore stable. The oldest captured group is the only nuance: if the capture did not reach the true end of history (final page full) and the dice-only count is not a positive multiple of 10, it may be an unfinished 10-pull continuing onto an uncaptured page; its captured prefix is still ordinal-stable, so it is exported and flagged `INCOMPLETE_TIMESTAMP_GROUP_EXPORTED` so the user knows to scroll further. If page 1 itself was not captured, the run falls back to the longest continuous block and emits `DID_NOT_START_AT_PAGE_1`.
Pages are anchored to the continuous run starting at page 1 (history always loads page 1 first), so the newest timestamp group's ordinal 0 is always captured. Ordinals are assigned in scan order (newest first), so ordinal 0 of a timestamp group is its newest record and any unseen continuation rows can only append after the captured ones with higher ordinals. Every exported UID is therefore stable, including a partially captured oldest 10-pull, so all decoded rows are exported. If page 1 itself was not captured, the run falls back to the longest continuous block and emits `DID_NOT_START_AT_PAGE_1`.
## Arc / Gashapon
@ -46,5 +46,5 @@ Pages are anchored to the continuous run starting at page 1 (history always load
- Pool: `Arc_MiracleBox`.
- Each response page normally contains 5 records.
- Arc timestamps use `unix_seconds = little_endian_u64(timestamp_raw) / 20000000 - 62135596800`.
- Arc pulls are treated as 10-pull timestamp groups. Like Monopoly, the oldest captured group is exported even if it is a 10-pull the scan stopped mid-way (its captured prefix is ordinal-stable) and flagged `INCOMPLETE_ARC_10_PULL_EXPORTED`.
- Arc pulls are treated as 10-pull timestamp groups. Like Monopoly, every captured group is exported, including the oldest one even if the scan stopped mid-10-pull (its captured prefix is ordinal-stable).
- Arc rows use the same `reward_type`, `reward_id`, `reward_name`, `reward_rank`, and `reward_key_hex` fields as Monopoly rows.

View file

@ -27,6 +27,6 @@
},
"notes": [
"Arc history is separate from Monopoly history.",
"Arc timestamp groups are 10-pull groups; an incomplete oldest group is still exported (stable prefix) with a warning."
"Arc timestamp groups are 10-pull groups; a partially captured oldest group is still exported (stable prefix)."
]
}

View file

@ -5,7 +5,6 @@ import json
from nte_history_exporter import console
from nte_history_exporter.adapters.mitmproxy_flows import decode_mitmproxy_flows
from nte_history_exporter.decoder.arc import arc_stability_warnings
from nte_history_exporter.decoder.boundary import annotate_groups
from nte_history_exporter.export.csv_export import write_csv
from nte_history_exporter.export.json_export import build_export_json
@ -39,13 +38,13 @@ def main(argv: list[str] | None = None) -> int:
decoded = decode_mitmproxy_flows(args.capture_source, args.flow_index)
if decoded["arc_rows"] and not decoded["rows"]:
rows = decoded["arc_rows"]
warnings = decoded["arc_warnings"] + arc_stability_warnings(rows)
warnings = decoded["arc_warnings"]
kind = "arc_miracle_box"
best_run = decoded["best_arc_run"]
pair_count = len(decoded["arc_pairs"])
else:
rows, group_warnings = annotate_groups(decoded["rows"])
warnings = decoded["run_warnings"] + group_warnings
rows = annotate_groups(decoded["rows"])
warnings = decoded["run_warnings"]
kind = decoded["best_run"][0][7] if decoded["best_run"] and len(decoded["best_run"][0]) > 7 else "permanent"
best_run = decoded["best_run"]
pair_count = len(decoded["pairs"])
@ -72,7 +71,6 @@ def main(argv: list[str] | None = None) -> int:
)
for warning in warnings:
console.print_warning(warning["code"], warning["reason"], warning.get("records"))
console.maybe_print_incomplete_hint(warnings)
print()
if args.debug:
console.print_note(f"CSV written: {out_path}")

View file

@ -119,37 +119,3 @@ def print_success(text: str) -> None:
def print_problem(text: str) -> None:
print(style(f" {text}", YELLOW, BOLD))
def print_boxed(lines: list[str], *codes: str) -> None:
inner = WIDTH - 4
border = " +" + "-" * (WIDTH - 2) + "+"
print(style(border, *codes))
for line in lines:
print(style(" | " + line.ljust(inner) + " |", *codes))
print(style(border, *codes))
# Warning codes for an oldest timestamp group that was exported despite possibly
# being an unfinished 10-pull. The exported rows are stable, so the warning is
# informational and the user can ignore it once the rest is captured.
INCOMPLETE_EXPORT_CODES = {
"INCOMPLETE_TIMESTAMP_GROUP_EXPORTED",
"INCOMPLETE_ARC_10_PULL_EXPORTED",
}
def maybe_print_incomplete_hint(warnings: list[dict]) -> None:
if not any(w.get("code") in INCOMPLETE_EXPORT_CODES for w in warnings):
return
print()
print_boxed(
[
"The INCOMPLETE_..._EXPORTED warning above is safe",
"to ignore if the partial 10-pull it captured is",
"already in your tracker. Otherwise restart",
"the capture and go to a further page",
],
GREEN,
BOLD,
)

View file

@ -141,59 +141,23 @@ def make_arc_uid(timestamp_raw: str, ordinal: int, arc_key_hex: str) -> str:
def annotate_arc_groups(rows: list[dict[str, Any]]) -> None:
# Pages are anchored at page 1, so ordinal 0 of every group is captured and
# all UIDs are stable -- even a partially captured oldest 10-pull, whose
# unseen rows can only append after the captured ones. All rows are exported.
groups: dict[str, list[int]] = defaultdict(list)
for index, row in enumerate(rows):
groups[row["timestamp_raw_hex"]].append(index)
group_items = list(groups.items())
for group_index, (timestamp_raw, indexes) in enumerate(group_items):
at_oldest_boundary = group_index == len(group_items) - 1
# Arc pulls are always 10-pulls, so the only group that can be short is
# the oldest one, where the scan stopped mid-10-pull. Its captured prefix
# is ordinal-stable (ordinal 0 is captured, unseen rows append after it),
# so export it and flag it rather than dropping it.
incomplete = at_oldest_boundary and len(indexes) % 10 != 0
skip_reason = (
"arc timestamp group is not a complete 10-pull in this capture; "
"exported rows are stable, scroll further to capture the rest"
if incomplete
else ""
)
for group_index, (timestamp_raw, indexes) in enumerate(groups.items()):
for ordinal, index in enumerate(indexes):
row = rows[index]
row["timestamp_group_index"] = group_index
row["timestamp_group_ordinal"] = ordinal
row["timestamp_group_size_seen"] = len(indexes)
row["uid"] = make_arc_uid(timestamp_raw, ordinal, row["reward_key_hex"])
row["uid_status"] = "incomplete_stable" if incomplete else "stable"
row["uid_status"] = "stable"
row["export_record"] = True
row["skip_reason"] = skip_reason
def arc_stability_warnings(rows: list[dict[str, Any]]) -> list[dict[str, Any]]:
warnings = []
seen = set()
for row in rows:
if row.get("uid_status") != "incomplete_stable":
continue
timestamp_raw = row["timestamp_raw_hex"]
if timestamp_raw in seen:
continue
seen.add(timestamp_raw)
group = [r for r in rows if r["timestamp_raw_hex"] == timestamp_raw]
warnings.append(
{
"code": "INCOMPLETE_ARC_10_PULL_EXPORTED",
"timestamp_raw": timestamp_raw,
"timestamp_decoded": row["timestamp_decoded"],
"records": len(group),
"reason": (
"arc timestamp group is not a complete 10-pull in this capture; "
"exported rows are stable, scroll further to capture the rest"
),
}
)
return warnings
row["skip_reason"] = ""
def select_continuous_arc_run(pairs: list[tuple]) -> tuple[list[tuple], list[dict[str, Any]]]:

View file

@ -100,23 +100,17 @@ def is_dice_record(row: dict[str, Any]) -> bool:
return False
def annotate_groups(rows: list[dict[str, Any]]) -> tuple[list[dict[str, Any]], list[dict[str, Any]]]:
def annotate_groups(rows: list[dict[str, Any]]) -> list[dict[str, Any]]:
"""Assign timestamp-group ordinals and stable UIDs to every decoded row.
Pages are anchored at page 1 (see select_continuous_run_from_page_1), so the
newest group's ordinal 0 is always captured and every UID is stable. The only
nuance is the oldest group: if the capture did not reach the true end of
history (final page full) and the dice-only count is not a complete multiple
of 10, it may be an unfinished 10-pull continuing onto an uncaptured page. Its
captured prefix is still ordinal-stable, so it is exported with an
informational warning rather than withheld.
newest group's ordinal 0 is always captured. Within a group, ordinal 0 is the
newest record and unseen continuation rows can only append after the captured
ones, so every exported UID is stable -- even a partially captured oldest
group. All decoded rows are therefore exported.
"""
if not rows:
return rows, []
last_page = max(int(r["page"]) for r in rows if str(r.get("page", "")).isdigit())
last_page_records = [r for r in rows if r.get("page") == last_page]
final_page_is_partial = len(last_page_records) < 5
return rows
groups: list[list[dict[str, Any]]] = []
current: list[dict[str, Any]] = []
@ -132,39 +126,18 @@ def annotate_groups(rows: list[dict[str, Any]]) -> tuple[list[dict[str, Any]], l
if current:
groups.append(current)
warnings: list[dict[str, Any]] = []
for group_index, group in enumerate(groups):
at_oldest_boundary = group_index == len(groups) - 1
dice_record_count = sum(1 for row in group if is_dice_record(row))
dice_complete = dice_record_count > 0 and dice_record_count % 10 == 0
incomplete = at_oldest_boundary and not final_page_is_partial and not dice_complete
skip_reason = (
"oldest timestamp group may continue onto the next uncaptured page; "
"exported rows are stable, scroll further to capture the rest"
if incomplete
else ""
)
if incomplete:
warnings.append(
{
"code": "INCOMPLETE_TIMESTAMP_GROUP_EXPORTED",
"timestamp_raw": group[0].get("timestamp_raw_hex", ""),
"timestamp_decoded": group[0].get("timestamp_decoded", ""),
"records": len(group),
"dice_records": dice_record_count,
"reason": skip_reason,
}
)
for ordinal, row in enumerate(group):
row["timestamp_group_index"] = group_index
row["timestamp_group_ordinal"] = ordinal
row["timestamp_group_size_seen"] = dice_record_count
row["timestamp_group_record_size_seen"] = len(group)
row["timestamp_group_boundary"] = "oldest" if at_oldest_boundary else ("newest" if group_index == 0 else "")
row["uid_status"] = "incomplete_stable" if incomplete else "stable"
row["timestamp_group_boundary"] = (
"oldest" if group_index == len(groups) - 1 else ("newest" if group_index == 0 else "")
)
row["uid_status"] = "stable"
row["uid"] = make_uid(row, ordinal)
row["export_record"] = True
row["skip_reason"] = skip_reason
return rows, warnings
row["skip_reason"] = ""
return rows

View file

@ -9,7 +9,6 @@ from pathlib import Path
from nte_history_exporter import console
from nte_history_exporter.constants import POOL_META
from nte_history_exporter.decoder.arc import arc_stability_warnings
from nte_history_exporter.decoder.boundary import annotate_groups, select_continuous_run_from_page_1
from nte_history_exporter.export.csv_export import write_csv
from nte_history_exporter.export.json_export import build_export_json
@ -88,11 +87,9 @@ def run_live_capture(
pairs = session.pairs_for_kind(kind)
best_run, run_warnings = select_continuous_run_from_page_1(pairs)
rows = session.build_rows(kind)
if kind == "arc_miracle_box":
warnings = run_warnings + arc_stability_warnings(rows)
else:
rows, group_warnings = annotate_groups(rows)
warnings = run_warnings + group_warnings
if kind != "arc_miracle_box":
rows = annotate_groups(rows)
warnings = run_warnings
pages_seen = [p[0] for p in best_run]
csv_path, json_path = export_paths(kind)
if write_debug_csv:
@ -123,7 +120,6 @@ def run_live_capture(
console.print_note("again, then reopen the history screen.")
return {"exports": []}
all_warnings = []
for item in exports:
scan = item["export"]["scan"]
console.print_export_summary(
@ -134,8 +130,6 @@ def run_live_capture(
)
for warning in scan["warnings"]:
console.print_warning(warning["code"], warning["reason"], warning.get("records"))
all_warnings.extend(scan["warnings"])
console.maybe_print_incomplete_hint(all_warnings)
print()
for item in exports:

View file

@ -20,7 +20,6 @@ from nte_history_exporter.constants import POOL_META
from nte_history_exporter.mappings import ARC_META, CHARACTERS, ITEMS, REWARDS_BY_ID
from nte_history_exporter.decoder.protocol import decode_reward_key, infer_reward_type
from nte_history_exporter.decoder.arc import (
arc_stability_warnings,
build_arc_rows_from_pairs,
decode_arc_key,
decode_arc_timestamp,
@ -163,19 +162,15 @@ class BoundaryExportTests(unittest.TestCase):
}
self.assertNotEqual(make_uid(row, 0), "5adcf52282e15445466863b271f3b745")
def test_pages_1_to_5_exports_dice_complete_oldest_group(self):
def test_pages_1_to_5_exports_every_row(self):
rows = load_reference_csv("monopoly_history_poc_13_pages_1_to_5_v4.csv")
annotated, warnings = annotate_groups(rows)
annotated = annotate_groups(rows)
exported = [row for row in annotated if row["export_record"] is True]
skipped = [row for row in annotated if row["export_record"] is False]
# The oldest group has exactly 10 dice rolls (plus a Points Gift), which
# proves the pull set is complete even though the scan stopped on a full page.
# Every decoded row is exported; boundary groups are never dropped.
self.assertEqual(len(annotated), 25)
self.assertEqual(len(exported), 25)
self.assertEqual(len(skipped), 0)
self.assertEqual(len(warnings), 0)
@staticmethod
def _synthetic_row(page, timestamp_hex, result_type):
@ -189,24 +184,21 @@ class BoundaryExportTests(unittest.TestCase):
"quantity": 1,
}
def test_oldest_group_with_partial_dice_count_exports_with_warning(self):
def test_oldest_group_with_partial_dice_count_exports_without_warning(self):
rows = [self._synthetic_row(1, "aa", "dice") for _ in range(5)]
rows += [self._synthetic_row(2, "bb", "dice") for _ in range(5)]
rows += [self._synthetic_row(3, "bb", "dice") for _ in range(4)]
rows += [self._synthetic_row(3, "bb", "points_gift")]
annotated, warnings = annotate_groups(rows)
annotated = annotate_groups(rows)
exported = [row for row in annotated if row["export_record"] is True]
# Oldest group is an unfinished 10-pull on a full final page. Its captured
# prefix is ordinal-stable, so it is exported with an informational warning
# rather than dropped.
# Oldest group is a partially captured 10-pull on a full final page. Its
# captured prefix is ordinal-stable, so it is exported with stable UIDs.
self.assertEqual(len(exported), 15)
self.assertEqual(len(warnings), 1)
self.assertEqual(warnings[0]["code"], "INCOMPLETE_TIMESTAMP_GROUP_EXPORTED")
self.assertEqual(warnings[0]["dice_records"], 9)
oldest = [row for row in annotated if row["timestamp_raw_hex"] == "bb"]
self.assertTrue(all(row["uid"] for row in oldest))
self.assertTrue(all(row["uid_status"] == "stable" for row in oldest))
self.assertEqual([row["timestamp_group_ordinal"] for row in oldest], list(range(10)))
def test_incomplete_oldest_prefix_keeps_stable_uids(self):
@ -215,8 +207,8 @@ class BoundaryExportTests(unittest.TestCase):
full += [self._synthetic_row(3, "bb", "dice") for _ in range(2)]
truncated = [r for r in full if r["page"] in (1, 2)]
full_rows, _ = annotate_groups([dict(r) for r in full])
trunc_rows, _ = annotate_groups([dict(r) for r in truncated])
full_rows = annotate_groups([dict(r) for r in full])
trunc_rows = annotate_groups([dict(r) for r in truncated])
full_uids = [r["uid"] for r in full_rows if r["timestamp_raw_hex"] == "bb"][:3]
trunc_uids = [r["uid"] for r in trunc_rows if r["timestamp_raw_hex"] == "bb"]
@ -230,11 +222,10 @@ class BoundaryExportTests(unittest.TestCase):
rows += [self._synthetic_row(2, "bb", "dice") for _ in range(5)]
rows += [self._synthetic_row(3, "bb", "dice") for _ in range(5)]
annotated, warnings = annotate_groups(rows)
annotated = annotate_groups(rows)
exported = [row for row in annotated if row["export_record"] is True]
self.assertEqual(len(exported), 15)
self.assertEqual(len(warnings), 0)
def test_run_selection_anchors_to_page_1_and_keeps_newest(self):
# Page 2's response was lost: captured pages 1, 3, 4, 5.
@ -257,7 +248,7 @@ class BoundaryExportTests(unittest.TestCase):
def test_full_reference_scan_exports_all_rows(self):
rows = load_reference_csv("monopoly_history_poc_10_all_44_pages_v4.csv")
annotated, warnings = annotate_groups(rows)
annotated = annotate_groups(rows)
exported = [row for row in annotated if row["export_record"] is True]
json_path = EXPORTS / "monopoly_history_export_10_all_44_pages_v4.json"
@ -270,12 +261,11 @@ class BoundaryExportTests(unittest.TestCase):
self.assertEqual(len(annotated), reference["scan"]["decoded_records"])
self.assertEqual(len(exported), reference["scan"]["exported_records"])
self.assertEqual(warnings, [])
def test_sanitized_export_omits_raw_packet_fields(self):
rows = load_reference_csv("monopoly_history_poc_13_pages_1_to_5_v4.csv")
annotated, warnings = annotate_groups(rows)
export = build_export_json(annotated, warnings)
annotated = annotate_groups(rows)
export = build_export_json(annotated, [])
self.assertEqual(export["format"], "nte-history-export")
self.assertIn("exporter", export)
@ -366,7 +356,7 @@ class BoundaryExportTests(unittest.TestCase):
self.assertEqual(decoded[0]["reward_type"], "arc")
self.assertEqual(decoded[0]["reward_key_hex"], reference_rows[0]["arc_key_hex"])
def test_arc_partial_timestamp_group_is_exported_with_warning(self):
def test_arc_partial_timestamp_group_is_exported_without_warning(self):
rows = load_arc_csv("arc_pages_1_to_5_v2.csv")
pairs = []
for page in range(1, 6):
@ -375,17 +365,14 @@ class BoundaryExportTests(unittest.TestCase):
pairs.append((page, page * 2, page, 1.0, page + 100, 1.1, response))
decoded = build_arc_rows_from_pairs(pairs)
exported = [row for row in decoded if row["export_record"] is True]
incomplete = [row for row in decoded if row["uid_status"] == "incomplete_stable"]
warnings = arc_stability_warnings(decoded)
# The oldest group is a 10-pull split by stopping at page 5 (5 of 10 rows).
# Its captured prefix is ordinal-stable, so it is exported, not dropped.
# Its captured prefix is ordinal-stable, so every row is exported with a
# stable UID and no warning.
self.assertEqual(len(decoded), 25)
self.assertEqual(len(exported), 25)
self.assertEqual(len(incomplete), 5)
self.assertEqual(len(warnings), 1)
self.assertEqual(warnings[0]["code"], "INCOMPLETE_ARC_10_PULL_EXPORTED")
self.assertTrue(all(row["uid"] for row in incomplete))
self.assertTrue(all(row["uid"] for row in decoded))
self.assertTrue(all(row["uid_status"] == "stable" for row in decoded))
def test_arc_incomplete_prefix_keeps_stable_uids(self):
rows = load_arc_csv("arc_pull_10_all_pages_v2.csv")
@ -468,13 +455,14 @@ class BoundaryExportTests(unittest.TestCase):
},
]
annotated, warnings = annotate_groups(rows)
annotated = annotate_groups(rows)
# Ordinals cover every row in the group, but the dice-only count drives
# timestamp_group_size_seen (2 dice in the 4-record group).
self.assertEqual([row["timestamp_group_ordinal"] for row in annotated[:4]], [0, 1, 2, 3])
self.assertEqual({row["timestamp_group_size_seen"] for row in annotated[:4]}, {2})
self.assertEqual({row["timestamp_group_record_size_seen"] for row in annotated[:4]}, {4})
self.assertEqual(warnings[0]["dice_records"], 1)
self.assertEqual(warnings[0]["records"], 1)
self.assertTrue(all(row["export_record"] for row in annotated))
def test_live_session_pairs_request_and_response(self):
session = LiveHistorySession("192.168.0.10")