Anchor scans to page 1; export stable prefixes

Anchor page-run selection to start at page 1 (select_continuous_run_from_page_1) instead of picking the longest run, so newest pages are preserved and the newest timestamp ordinal 0 is always captured. Change boundary logic to always assign stable UIDs and export the captured prefix of an oldest timestamp group even when it may be an unfinished 10-pull, emitting informational warnings (INCOMPLETE_TIMESTAMP_GROUP_EXPORTED / INCOMPLETE_ARC_10_PULL_EXPORTED) instead of dropping rows. Apply the same policy to Arc groups and update arc group annotation and warnings accordingly. Add a new console module for improved CLI output and update cli, live_capture, adapters, and session code to use the new run selection and console helpers. Update docs, mappings, and tests to reflect the new boundary/export behavior and messaging.
This commit is contained in:
Golumpa 2026-06-11 12:22:34 +01:00
parent 0a1a19939b
commit 20b2086299
14 changed files with 402 additions and 200 deletions

View file

@ -32,7 +32,7 @@ Live capture, Windows:
.\run-exporter.ps1 --live .\run-exporter.ps1 --live
``` ```
Launch it before logging in, then open any supported history board in game. The tool keeps listening until you press any key. Exports are written under `exports\` as `Permanent_<date_time>.json`, `Limited_<date_time>.json`, or `Arc_<date_time>.json`. Launch it before pressing Start on the game's main menu so the game's UDP connection can be captured, then open any supported history board in game. If you are already in game, log out to the main menu and enter again. The tool keeps listening until you press any key. Exports are written under `exports\` as `Permanent_<date_time>.json`, `Limited_<date_time>.json`, or `Arc_<date_time>.json`.
Pass `--debug` to also write the full research CSV next to each JSON export. Pass `--debug` to also write the full research CSV next to each JSON export.
@ -55,9 +55,11 @@ Live capture needs a local admin-capable packet socket on Windows. The prototype
NTE history records do not appear to contain a unique server-side roll ID. UIDs are generated from decoded record fields and the record's order within all rows sharing the same raw timestamp. NTE history records do not appear to contain a unique server-side roll ID. UIDs are generated from decoded record fields and the record's order within all rows sharing the same raw timestamp.
Because 10-pulls can span page boundaries, partial timestamp groups can produce unstable UIDs. Normal mode exports only complete/stable timestamp groups. Boundary groups are skipped with warnings when the exporter cannot prove they are complete. History always loads page 1 first and is scrolled downward, so the exporter anchors to the continuous run of pages starting at page 1 and ignores anything after the first gap (with a warning). This keeps the newest pages even if a later page is lost, and guarantees the newest timestamp group's ordinal 0 is captured.
For Monopoly, Points Gift and Chase Reward rows stay in the timestamp group for UID ordinal generation, but only `result_type = dice` rows count toward pull-set sizing. An oldest boundary group whose dice-roll count is a complete multiple of 10 is treated as a finished pull set and exported even when the scan stopped on a full page. Arc groups are expected to be complete 10-pull timestamp groups. Within a timestamp group, ordinal 0 is the newest record and unseen rows can only append after the captured ones, so every exported UID is stable. The one nuance is the oldest captured group: if the capture did not reach the true end of history and its dice-roll count is not a complete multiple of 10, it may be an unfinished 10-pull continuing onto an uncaptured page. Its captured prefix is still ordinal-stable, so it is exported with an informational warning telling you to scroll further to capture the rest.
For Monopoly, Points Gift and Chase Reward rows stay in the timestamp group for UID ordinal generation, but only `result_type = dice` rows count toward pull-set sizing. Arc pulls are always 10-pulls, so the same rule applies with a fixed group size of 10. In both systems the oldest captured group is exported even if it is an unfinished pull set (its captured prefix is ordinal-stable), and flagged with a warning so you know to scroll further on a later scan.
## Current Adapters ## Current Adapters

View file

@ -20,7 +20,7 @@ The sanitized JSON export uses:
}, },
"scan": { "scan": {
"mode": "stable_only", "mode": "stable_only",
"boundary_policy": "drop_incomplete_timestamp_groups", "boundary_policy": "export_ordinal_stable_groups",
"decoded_records": 0, "decoded_records": 0,
"exported_records": 0, "exported_records": 0,
"skipped_records": 0, "skipped_records": 0,

View file

@ -7,8 +7,8 @@ Known limitations:
- Other NTE banners are not implemented yet. - Other NTE banners are not implemented yet.
- The game appears not to provide a unique server-side roll ID in the decoded record body. - The game appears not to provide a unique server-side roll ID in the decoded record body.
- UIDs are generated deterministically from decoded fields and timestamp-group order. - UIDs are generated deterministically from decoded fields and timestamp-group order.
- Boundary timestamp groups may be skipped to avoid exporting unstable data. - An incomplete oldest timestamp group is still exported (its captured prefix has stable UIDs) and flagged so the user can scroll further on a later scan.
- Page gaps are ignored outside the longest continuous run and reported as warnings. - Pages are anchored to the continuous run starting at page 1; pages after the first gap are ignored and reported as warnings.
- Live capture currently uses Windows raw sockets and requires administrator permission. - Live capture currently uses Windows raw sockets and requires administrator permission.
- The file adapter reads mitmproxy `.flows` captures for research and testing. - The file adapter reads mitmproxy `.flows` captures for research and testing.
- The more stable and reliable Npcap/libpcap capture is not implemented yet. - The more stable and reliable Npcap/libpcap capture is not implemented yet.

View file

@ -36,7 +36,7 @@ Page and row numbers are research metadata only. They must not be used for perma
Timestamp groups keep all records with the same raw timestamp together for UID ordinal generation. For boundary/group-size detection, only `result_type = dice` rows count as pull-set members; Points Gift and Chase Reward rows stay in the group but do not increase the dice-only group count. Timestamp groups keep all records with the same raw timestamp together for UID ordinal generation. For boundary/group-size detection, only `result_type = dice` rows count as pull-set members; Points Gift and Chase Reward rows stay in the group but do not increase the dice-only group count.
Oldest-boundary groups are exported when their dice-only count is a positive multiple of 10: the pull set is provably finished, and any unseen same-timestamp continuation rows can only be non-dice tails that sort after the seen rows, so exported ordinals and UIDs stay stable. Oldest groups with other dice counts are still dropped unless the capture reached the true end of history. Newest-boundary groups remain positional-only: when a scan starts mid-history, unseen newer rows would shift ordinals, so dice count is never accepted as proof there. Pages are anchored to the continuous run starting at page 1 (history always loads page 1 first), so the newest timestamp group's ordinal 0 is always captured. Ordinals are assigned in scan order (newest first), so ordinal 0 of a timestamp group is its newest record and any unseen continuation rows can only append after the captured ones with higher ordinals. Every exported UID is therefore stable. The oldest captured group is the only nuance: if the capture did not reach the true end of history (final page full) and the dice-only count is not a positive multiple of 10, it may be an unfinished 10-pull continuing onto an uncaptured page; its captured prefix is still ordinal-stable, so it is exported and flagged `INCOMPLETE_TIMESTAMP_GROUP_EXPORTED` so the user knows to scroll further. If page 1 itself was not captured, the run falls back to the longest continuous block and emits `DID_NOT_START_AT_PAGE_1`.
## Arc / Gashapon ## Arc / Gashapon
@ -46,5 +46,5 @@ Oldest-boundary groups are exported when their dice-only count is a positive mul
- Pool: `Arc_MiracleBox`. - Pool: `Arc_MiracleBox`.
- Each response page normally contains 5 records. - Each response page normally contains 5 records.
- Arc timestamps use `unix_seconds = little_endian_u64(timestamp_raw) / 20000000 - 62135596800`. - Arc timestamps use `unix_seconds = little_endian_u64(timestamp_raw) / 20000000 - 62135596800`.
- Arc pulls are treated as 10-pull timestamp groups; incomplete groups are skipped by default. - Arc pulls are treated as 10-pull timestamp groups. Like Monopoly, the oldest captured group is exported even if it is a 10-pull the scan stopped mid-way (its captured prefix is ordinal-stable) and flagged `INCOMPLETE_ARC_10_PULL_EXPORTED`.
- Arc rows use the same `reward_type`, `reward_id`, `reward_name`, `reward_rank`, and `reward_key_hex` fields as Monopoly rows. - Arc rows use the same `reward_type`, `reward_id`, `reward_name`, `reward_rank`, and `reward_key_hex` fields as Monopoly rows.

View file

@ -27,6 +27,6 @@
}, },
"notes": [ "notes": [
"Arc history is separate from Monopoly history.", "Arc history is separate from Monopoly history.",
"Arc timestamp groups are expected to be complete 10-pull groups." "Arc timestamp groups are 10-pull groups; an incomplete oldest group is still exported (stable prefix) with a warning."
] ]
} }

View file

@ -4,7 +4,7 @@ import struct
from pathlib import Path from pathlib import Path
from typing import Any from typing import Any
from nte_history_exporter.decoder.boundary import longest_monotonic_page_run from nte_history_exporter.decoder.boundary import select_continuous_run_from_page_1
from nte_history_exporter.decoder.arc import ( from nte_history_exporter.decoder.arc import (
arc_request_page, arc_request_page,
build_arc_rows_from_pairs, build_arc_rows_from_pairs,
@ -128,7 +128,7 @@ def decode_mitmproxy_flows(path: str | Path, flow_index: int | None = None) -> d
if response_index is not None: if response_index is not None:
pairs.append((page, offset, i, ts, response_index, response_ts, response_content, kind)) pairs.append((page, offset, i, ts, response_index, response_ts, response_content, kind))
best_run = longest_monotonic_page_run(pairs) best_run, run_warnings = select_continuous_run_from_page_1(pairs)
rows_out = build_rows_from_pairs(best_run) rows_out = build_rows_from_pairs(best_run)
best_arc_run, arc_warnings = select_continuous_arc_run(arc_pairs) best_arc_run, arc_warnings = select_continuous_arc_run(arc_pairs)
arc_rows = build_arc_rows_from_pairs(best_arc_run) arc_rows = build_arc_rows_from_pairs(best_arc_run)
@ -137,10 +137,10 @@ def decode_mitmproxy_flows(path: str | Path, flow_index: int | None = None) -> d
"flow_index": resolved_flow_index, "flow_index": resolved_flow_index,
"pairs": pairs, "pairs": pairs,
"best_run": best_run, "best_run": best_run,
"run_warnings": run_warnings,
"rows": rows_out, "rows": rows_out,
"arc_pairs": arc_pairs, "arc_pairs": arc_pairs,
"best_arc_run": best_arc_run, "best_arc_run": best_arc_run,
"arc_rows": arc_rows, "arc_rows": arc_rows,
"arc_warnings": arc_warnings, "arc_warnings": arc_warnings,
"starts_from_page_1": bool(best_run and best_run[0][0] == 1),
} }

View file

@ -3,9 +3,10 @@ from __future__ import annotations
import argparse import argparse
import json import json
from nte_history_exporter import console
from nte_history_exporter.adapters.mitmproxy_flows import decode_mitmproxy_flows from nte_history_exporter.adapters.mitmproxy_flows import decode_mitmproxy_flows
from nte_history_exporter.decoder.arc import arc_stability_warnings from nte_history_exporter.decoder.arc import arc_stability_warnings
from nte_history_exporter.decoder.boundary import annotate_groups, page_gap_warnings from nte_history_exporter.decoder.boundary import annotate_groups
from nte_history_exporter.export.csv_export import write_csv from nte_history_exporter.export.csv_export import write_csv
from nte_history_exporter.export.json_export import build_export_json from nte_history_exporter.export.json_export import build_export_json
from nte_history_exporter.live_capture.runner import export_paths, run_live_capture from nte_history_exporter.live_capture.runner import export_paths, run_live_capture
@ -21,35 +22,18 @@ def build_parser() -> argparse.ArgumentParser:
parser.add_argument("--interface-ip", default=None, help="local IPv4 address to bind for live capture") parser.add_argument("--interface-ip", default=None, help="local IPv4 address to bind for live capture")
parser.add_argument("--no-clipboard", action="store_true", help="do not copy live exports to clipboard") parser.add_argument("--no-clipboard", action="store_true", help="do not copy live exports to clipboard")
parser.add_argument("--debug", action="store_true", help="also write research CSVs next to the JSON exports") parser.add_argument("--debug", action="store_true", help="also write research CSVs next to the JSON exports")
parser.add_argument(
"--allow-boundary-records",
action="store_true",
help="Debug only: export boundary groups even when they may be partial.",
)
return parser return parser
def main(argv: list[str] | None = None) -> int: def main(argv: list[str] | None = None) -> int:
args = build_parser().parse_args(argv) args = build_parser().parse_args(argv)
console.print_banner()
if args.live or not args.capture_source: if args.live or not args.capture_source:
result = run_live_capture( run_live_capture(
interface_ip=args.interface_ip, interface_ip=args.interface_ip,
copy_clipboard=not args.no_clipboard, copy_clipboard=not args.no_clipboard,
write_debug_csv=args.debug, write_debug_csv=args.debug,
) )
for item in result["exports"]:
export = item["export"]
print(
"{banner}: decoded {decoded}, exported {exported}, skipped {skipped}.".format(
banner=export["banner"]["name"],
decoded=export["scan"]["decoded_records"],
exported=export["scan"]["exported_records"],
skipped=export["scan"]["skipped_records"],
)
)
for warning in export["scan"]["warnings"]:
suffix = f" ({warning['records']} records)" if "records" in warning else ""
print(f"WARNING {warning['code']}: {warning['reason']}{suffix}")
return 0 return 0
decoded = decode_mitmproxy_flows(args.capture_source, args.flow_index) decoded = decode_mitmproxy_flows(args.capture_source, args.flow_index)
@ -60,12 +44,8 @@ def main(argv: list[str] | None = None) -> int:
best_run = decoded["best_arc_run"] best_run = decoded["best_arc_run"]
pair_count = len(decoded["arc_pairs"]) pair_count = len(decoded["arc_pairs"])
else: else:
rows, warnings = annotate_groups( rows, group_warnings = annotate_groups(decoded["rows"])
decoded["rows"], warnings = decoded["run_warnings"] + group_warnings
starts_from_page_1=decoded["starts_from_page_1"],
stable_only=not args.allow_boundary_records,
)
warnings = page_gap_warnings(decoded["pairs"], decoded["best_run"]) + warnings
kind = decoded["best_run"][0][7] if decoded["best_run"] and len(decoded["best_run"][0]) > 7 else "permanent" kind = decoded["best_run"][0][7] if decoded["best_run"] and len(decoded["best_run"][0]) > 7 else "permanent"
best_run = decoded["best_run"] best_run = decoded["best_run"]
pair_count = len(decoded["pairs"]) pair_count = len(decoded["pairs"])
@ -83,19 +63,20 @@ def main(argv: list[str] | None = None) -> int:
) )
json_path.write_text(json.dumps(export, ensure_ascii=False, indent=2), encoding="utf-8") json_path.write_text(json.dumps(export, ensure_ascii=False, indent=2), encoding="utf-8")
if args.debug: console.print_results_header()
print(f"CSV written: {out_path}") console.print_export_summary(
print(f"JSON written: {json_path}") export["banner"]["name"],
print( export["scan"]["decoded_records"],
"Decoded {decoded}, exported {exported}, skipped {skipped}.".format( export["scan"]["exported_records"],
decoded=export["scan"]["decoded_records"], export["scan"]["skipped_records"],
exported=export["scan"]["exported_records"],
skipped=export["scan"]["skipped_records"],
)
) )
for warning in warnings: for warning in warnings:
suffix = f" ({warning['records']} records)" if "records" in warning else "" console.print_warning(warning["code"], warning["reason"], warning.get("records"))
print(f"WARNING {warning['code']}: {warning['reason']}{suffix}") console.maybe_print_incomplete_hint(warnings)
print()
if args.debug:
console.print_note(f"CSV written: {out_path}")
console.print_note(f"Export written: {json_path}")
return 0 return 0

View file

@ -0,0 +1,155 @@
from __future__ import annotations
import os
import sys
from nte_history_exporter.constants import EXPORTER_VERSION, GAME_NAME
WIDTH = 58
RESET = "\x1b[0m"
BOLD = "\x1b[1m"
DIM = "\x1b[2m"
CYAN = "\x1b[36m"
GREEN = "\x1b[32m"
YELLOW = "\x1b[33m"
_ansi: bool | None = None
def _enable_ansi() -> bool:
if not hasattr(sys.stdout, "isatty") or not sys.stdout.isatty():
return False
if os.name != "nt":
return True
try:
import ctypes
kernel32 = ctypes.windll.kernel32
handle = kernel32.GetStdHandle(-11)
mode = ctypes.c_uint32()
if not kernel32.GetConsoleMode(handle, ctypes.byref(mode)):
return False
ENABLE_VIRTUAL_TERMINAL_PROCESSING = 0x0004
return bool(kernel32.SetConsoleMode(handle, mode.value | ENABLE_VIRTUAL_TERMINAL_PROCESSING))
except Exception:
return False
def ansi_enabled() -> bool:
global _ansi
if _ansi is None:
_ansi = _enable_ansi()
return _ansi
def style(text: str, *codes: str) -> str:
if not codes or not ansi_enabled():
return text
return "".join(codes) + text + RESET
def rule(char: str = "-") -> str:
return style(char * WIDTH, DIM)
def print_banner() -> None:
print()
print(rule("="))
print(style(f" NTE History Exporter v{EXPORTER_VERSION}", BOLD, CYAN))
print(style(f" {GAME_NAME} pull history -> tracker JSON", DIM))
print(rule("="))
def print_live_instructions(local_ip: str) -> None:
print()
print(style(" Listening on ", DIM) + style(local_ip, BOLD))
print()
print(style(" How to export your pull history", BOLD))
print(" 1. This tool must be running BEFORE you press Start on")
print(" the game's main menu, or the game connection cannot")
print(" be captured. Already in game? Log out to the main")
print(" menu and enter again.")
print(" 2. Open a supported history screen:")
print(style(" Monopoly > Standard Board history", CYAN))
print(style(" Monopoly > Limited Character Board history", CYAN))
print(style(" Gashapon > Arc Miracle Box history", CYAN))
print(" 3. Start at page 1 and scroll down through every page")
print(" you want exported.")
print(" 4. Scroll one page past where you plan to stop so the")
print(" last pull group can be confirmed as complete.")
print(" 5. You can open several boards in the same session.")
print()
print(style(" Waiting for history pages... press any key here when done.", BOLD, GREEN))
print(rule())
def print_page_captured(label: str, page: int | None) -> None:
print(style(" + ", GREEN, BOLD) + label + style(f" page {page}", DIM))
def print_results_header() -> None:
print()
print(rule())
print(style(" Results", BOLD))
print(rule())
def print_export_summary(name: str, decoded: int, exported: int, skipped: int) -> None:
counts = [
style(f"decoded {decoded}", DIM),
style(f"exported {exported}", GREEN, BOLD),
style(f"skipped {skipped}", YELLOW if skipped else DIM),
]
print(f" {name:<30}" + " ".join(counts))
def print_warning(code: str, reason: str, records: int | None = None) -> None:
suffix = f" ({records} records)" if records is not None else ""
print(style(f" ! {code}: {reason}{suffix}", YELLOW))
def print_note(text: str) -> None:
print(style(f" {text}", DIM))
def print_success(text: str) -> None:
print(style(f" {text}", GREEN, BOLD))
def print_problem(text: str) -> None:
print(style(f" {text}", YELLOW, BOLD))
def print_boxed(lines: list[str], *codes: str) -> None:
inner = WIDTH - 4
border = " +" + "-" * (WIDTH - 2) + "+"
print(style(border, *codes))
for line in lines:
print(style(" | " + line.ljust(inner) + " |", *codes))
print(style(border, *codes))
# Warning codes for an oldest timestamp group that was exported despite possibly
# being an unfinished 10-pull. The exported rows are stable, so the warning is
# informational and the user can ignore it once the rest is captured.
INCOMPLETE_EXPORT_CODES = {
"INCOMPLETE_TIMESTAMP_GROUP_EXPORTED",
"INCOMPLETE_ARC_10_PULL_EXPORTED",
}
def maybe_print_incomplete_hint(warnings: list[dict]) -> None:
if not any(w.get("code") in INCOMPLETE_EXPORT_CODES for w in warnings):
return
print()
print_boxed(
[
"The INCOMPLETE_..._EXPORTED warning above is safe",
"to ignore if the partial 10-pull it captured is",
"already in your tracker. Otherwise restart",
"the capture and go to a further page",
],
GREEN,
BOLD,
)

View file

@ -19,7 +19,7 @@ from nte_history_exporter.constants import (
GAME_UID_PART, GAME_UID_PART,
POOL_META, POOL_META,
) )
from nte_history_exporter.decoder.boundary import longest_monotonic_page_run from nte_history_exporter.decoder.boundary import select_continuous_run_from_page_1
from nte_history_exporter.decoder.run import fmt_packet_time from nte_history_exporter.decoder.run import fmt_packet_time
from nte_history_exporter.mappings import ARC_META from nte_history_exporter.mappings import ARC_META
@ -145,24 +145,36 @@ def annotate_arc_groups(rows: list[dict[str, Any]]) -> None:
for index, row in enumerate(rows): for index, row in enumerate(rows):
groups[row["timestamp_raw_hex"]].append(index) groups[row["timestamp_raw_hex"]].append(index)
for group_index, (timestamp_raw, indexes) in enumerate(groups.items()): group_items = list(groups.items())
complete = len(indexes) % 10 == 0 for group_index, (timestamp_raw, indexes) in enumerate(group_items):
at_oldest_boundary = group_index == len(group_items) - 1
# Arc pulls are always 10-pulls, so the only group that can be short is
# the oldest one, where the scan stopped mid-10-pull. Its captured prefix
# is ordinal-stable (ordinal 0 is captured, unseen rows append after it),
# so export it and flag it rather than dropping it.
incomplete = at_oldest_boundary and len(indexes) % 10 != 0
skip_reason = (
"arc timestamp group is not a complete 10-pull in this capture; "
"exported rows are stable, scroll further to capture the rest"
if incomplete
else ""
)
for ordinal, index in enumerate(indexes): for ordinal, index in enumerate(indexes):
row = rows[index] row = rows[index]
row["timestamp_group_index"] = group_index row["timestamp_group_index"] = group_index
row["timestamp_group_ordinal"] = ordinal row["timestamp_group_ordinal"] = ordinal
row["timestamp_group_size_seen"] = len(indexes) row["timestamp_group_size_seen"] = len(indexes)
row["uid"] = make_arc_uid(timestamp_raw, ordinal, row["reward_key_hex"]) row["uid"] = make_arc_uid(timestamp_raw, ordinal, row["reward_key_hex"])
row["uid_status"] = "stable" if complete else "skipped_incomplete_timestamp_group" row["uid_status"] = "incomplete_stable" if incomplete else "stable"
row["export_record"] = complete row["export_record"] = True
row["skip_reason"] = "" if complete else "arc timestamp group is not a complete 10-pull in this capture" row["skip_reason"] = skip_reason
def arc_stability_warnings(rows: list[dict[str, Any]]) -> list[dict[str, Any]]: def arc_stability_warnings(rows: list[dict[str, Any]]) -> list[dict[str, Any]]:
warnings = [] warnings = []
seen = set() seen = set()
for row in rows: for row in rows:
if row.get("export_record") is True: if row.get("uid_status") != "incomplete_stable":
continue continue
timestamp_raw = row["timestamp_raw_hex"] timestamp_raw = row["timestamp_raw_hex"]
if timestamp_raw in seen: if timestamp_raw in seen:
@ -171,38 +183,18 @@ def arc_stability_warnings(rows: list[dict[str, Any]]) -> list[dict[str, Any]]:
group = [r for r in rows if r["timestamp_raw_hex"] == timestamp_raw] group = [r for r in rows if r["timestamp_raw_hex"] == timestamp_raw]
warnings.append( warnings.append(
{ {
"code": "INCOMPLETE_ARC_10_PULL_DROPPED", "code": "INCOMPLETE_ARC_10_PULL_EXPORTED",
"timestamp_raw": timestamp_raw, "timestamp_raw": timestamp_raw,
"timestamp_decoded": row["timestamp_decoded"], "timestamp_decoded": row["timestamp_decoded"],
"records": len(group), "records": len(group),
"reason": "arc timestamp group is not a complete 10-pull in this capture", "reason": (
"arc timestamp group is not a complete 10-pull in this capture; "
"exported rows are stable, scroll further to capture the rest"
),
} }
) )
return warnings return warnings
def select_continuous_arc_run(pairs: list[tuple]) -> tuple[list[tuple], list[dict[str, Any]]]: def select_continuous_arc_run(pairs: list[tuple]) -> tuple[list[tuple], list[dict[str, Any]]]:
warnings: list[dict[str, Any]] = [] return select_continuous_run_from_page_1(pairs)
if not pairs:
return [], warnings
pairs_by_page = {pair[0]: pair for pair in pairs}
seen_pages = sorted(pairs_by_page)
if 1 in pairs_by_page:
selected_pages = []
page = 1
while page in pairs_by_page:
selected_pages.append(page)
page += 1
if len(selected_pages) < len(seen_pages):
ignored = [page for page in seen_pages if page not in selected_pages]
warnings.append(
{
"code": "PAGE_GAP_DETECTED",
"ignored_pages": ignored,
"reason": f"Using continuous pages 1-{selected_pages[-1]}; ignored later pages {ignored}.",
}
)
return [pairs_by_page[page] for page in selected_pages], warnings
warnings.append({"code": "DID_NOT_START_AT_PAGE_1", "reason": "Arc history scan did not start at page 1."})
return longest_monotonic_page_run(pairs), warnings

View file

@ -40,31 +40,50 @@ def longest_monotonic_page_run(pairs: list[tuple]) -> list[tuple]:
return max(runs, key=len) if runs else [] return max(runs, key=len) if runs else []
def page_gap_warnings(pairs: list[tuple], best_run: list[tuple]) -> list[dict[str, Any]]: def select_continuous_run_from_page_1(pairs: list[tuple]) -> tuple[list[tuple], list[dict[str, Any]]]:
warnings: list[dict[str, Any]] = [] """Pick the run of pages starting at page 1 and continuing without gaps.
if len(pairs) < 2:
return warnings
best_run_ids = {id(pair) for pair in best_run} History always loads page 1 first and is scrolled downward, so the page-1
ignored_pages = [pair[0] for pair in pairs if id(pair) not in best_run_ids] run holds the newest, contiguous history. Anchoring here (instead of the
previous_page = pairs[0][0] longest run anywhere) keeps the newest pages even when a later packet is
for pair in pairs[1:]: lost, and guarantees the newest timestamp group's ordinal 0 is captured so
page = pair[0] every exported UID is stable. Pages after the first gap are ignored with a
if page != previous_page + 1: warning. If page 1 itself was not captured we fall back to the longest run
warning = { and warn that the result may be unstable.
"""
warnings: list[dict[str, Any]] = []
if not pairs:
return [], warnings
pairs_by_page = {pair[0]: pair for pair in pairs}
seen_pages = sorted(pairs_by_page)
if 1 in pairs_by_page:
selected_pages: list[int] = []
page = 1
while page in pairs_by_page:
selected_pages.append(page)
page += 1
if len(selected_pages) < len(seen_pages):
ignored = [p for p in seen_pages if p not in selected_pages]
warnings.append(
{
"code": "PAGE_GAP_DETECTED", "code": "PAGE_GAP_DETECTED",
"previous_page": previous_page, "ignored_pages": ignored,
"next_page": page,
"ignored_pages": ignored_pages,
"reason": ( "reason": (
f"Page gap detected: saw page {previous_page} then page {page}. " f"Page gap detected after page {selected_pages[-1]}; "
"Pages outside the longest continuous run were ignored for stable dedupe. " f"ignored later pages {ignored}. Re-scan or scroll more slowly."
"Re-scan or scroll more slowly."
), ),
} }
warnings.append(warning) )
previous_page = page return [pairs_by_page[p] for p in selected_pages], warnings
return warnings
warnings.append(
{
"code": "DID_NOT_START_AT_PAGE_1",
"reason": "Page 1 was not captured; results may be unstable. Re-scan from the top.",
}
)
return longest_monotonic_page_run(pairs), warnings
def is_dice_record(row: dict[str, Any]) -> bool: def is_dice_record(row: dict[str, Any]) -> bool:
@ -81,12 +100,17 @@ def is_dice_record(row: dict[str, Any]) -> bool:
return False return False
def annotate_groups( def annotate_groups(rows: list[dict[str, Any]]) -> tuple[list[dict[str, Any]], list[dict[str, Any]]]:
rows: list[dict[str, Any]], """Assign timestamp-group ordinals and stable UIDs to every decoded row.
*,
starts_from_page_1: bool = True, Pages are anchored at page 1 (see select_continuous_run_from_page_1), so the
stable_only: bool = True, newest group's ordinal 0 is always captured and every UID is stable. The only
) -> tuple[list[dict[str, Any]], list[dict[str, Any]]]: nuance is the oldest group: if the capture did not reach the true end of
history (final page full) and the dice-only count is not a complete multiple
of 10, it may be an unfinished 10-pull continuing onto an uncaptured page. Its
captured prefix is still ordinal-stable, so it is exported with an
informational warning rather than withheld.
"""
if not rows: if not rows:
return rows, [] return rows, []
@ -110,28 +134,21 @@ def annotate_groups(
warnings: list[dict[str, Any]] = [] warnings: list[dict[str, Any]] = []
for group_index, group in enumerate(groups): for group_index, group in enumerate(groups):
at_newest_boundary = group_index == 0
at_oldest_boundary = group_index == len(groups) - 1 at_oldest_boundary = group_index == len(groups) - 1
dice_records_in_group = [row for row in group if is_dice_record(row)] dice_record_count = sum(1 for row in group if is_dice_record(row))
dice_record_count = len(dice_records_in_group)
# A complete multiple of 10 dice rolls proves the pull set is finished.
# Unseen continuation rows can only be non-dice tails that sort after the
# seen rows, so exported ordinals (and UIDs) stay stable.
dice_complete = dice_record_count > 0 and dice_record_count % 10 == 0 dice_complete = dice_record_count > 0 and dice_record_count % 10 == 0
group_status = "stable" incomplete = at_oldest_boundary and not final_page_is_partial and not dice_complete
skip_reason = "" skip_reason = (
"oldest timestamp group may continue onto the next uncaptured page; "
"exported rows are stable, scroll further to capture the rest"
if incomplete
else ""
)
if at_newest_boundary and not starts_from_page_1: if incomplete:
group_status = "dropped_boundary_group"
skip_reason = "newest timestamp group may be partial because scan did not start from page 1"
elif at_oldest_boundary and not final_page_is_partial and not dice_complete:
group_status = "dropped_boundary_group"
skip_reason = "oldest timestamp group may continue onto the next uncaptured page"
if group_status != "stable":
warnings.append( warnings.append(
{ {
"code": "PARTIAL_TIMESTAMP_GROUP_DROPPED", "code": "INCOMPLETE_TIMESTAMP_GROUP_EXPORTED",
"timestamp_raw": group[0].get("timestamp_raw_hex", ""), "timestamp_raw": group[0].get("timestamp_raw_hex", ""),
"timestamp_decoded": group[0].get("timestamp_decoded", ""), "timestamp_decoded": group[0].get("timestamp_decoded", ""),
"records": len(group), "records": len(group),
@ -145,19 +162,9 @@ def annotate_groups(
row["timestamp_group_ordinal"] = ordinal row["timestamp_group_ordinal"] = ordinal
row["timestamp_group_size_seen"] = dice_record_count row["timestamp_group_size_seen"] = dice_record_count
row["timestamp_group_record_size_seen"] = len(group) row["timestamp_group_record_size_seen"] = len(group)
row["timestamp_group_boundary"] = ",".join( row["timestamp_group_boundary"] = "oldest" if at_oldest_boundary else ("newest" if group_index == 0 else "")
name row["uid_status"] = "incomplete_stable" if incomplete else "stable"
for name, yes in [("newest", at_newest_boundary), ("oldest", at_oldest_boundary)]
if yes
)
if group_status == "stable":
row["uid_status"] = "stable"
row["uid"] = make_uid(row, ordinal) row["uid"] = make_uid(row, ordinal)
row["export_record"] = True row["export_record"] = True
row["skip_reason"] = ""
else:
row["uid_status"] = group_status
row["uid"] = "" if stable_only else make_uid(row, ordinal)
row["export_record"] = not stable_only
row["skip_reason"] = skip_reason row["skip_reason"] = skip_reason
return rows, warnings return rows, warnings

View file

@ -27,7 +27,7 @@ def build_export_json(
pool = next((meta for meta in POOL_META.values() if meta["id"] == pool_group_id), POOL_META["permanent"]) pool = next((meta for meta in POOL_META.values() if meta["id"] == pool_group_id), POOL_META["permanent"])
scan: dict[str, Any] = { scan: dict[str, Any] = {
"mode": "stable_only", "mode": "stable_only",
"boundary_policy": "drop_incomplete_timestamp_groups", "boundary_policy": "export_ordinal_stable_groups",
"decoded_records": len(rows), "decoded_records": len(rows),
"exported_records": len(exported), "exported_records": len(exported),
"skipped_records": len(rows) - len(exported), "skipped_records": len(rows) - len(exported),

View file

@ -7,9 +7,10 @@ import time
from datetime import datetime from datetime import datetime
from pathlib import Path from pathlib import Path
from nte_history_exporter import console
from nte_history_exporter.constants import POOL_META from nte_history_exporter.constants import POOL_META
from nte_history_exporter.decoder.arc import arc_stability_warnings, select_continuous_arc_run from nte_history_exporter.decoder.arc import arc_stability_warnings
from nte_history_exporter.decoder.boundary import annotate_groups, page_gap_warnings from nte_history_exporter.decoder.boundary import annotate_groups, select_continuous_run_from_page_1
from nte_history_exporter.export.csv_export import write_csv from nte_history_exporter.export.csv_export import write_csv
from nte_history_exporter.export.json_export import build_export_json from nte_history_exporter.export.json_export import build_export_json
from nte_history_exporter.live_capture.session import LiveHistorySession, UdpPacket from nte_history_exporter.live_capture.session import LiveHistorySession, UdpPacket
@ -52,8 +53,7 @@ def run_live_capture(
session = LiveHistorySession(local_ip) session = LiveHistorySession(local_ip)
sock = open_raw_udp_socket(local_ip) sock = open_raw_udp_socket(local_ip)
print(f"Device 0 ready~! Listening on {local_ip}") console.print_live_instructions(local_ip)
print("Open either Monopoly history board. Capture will stay open until you press any key.")
try: try:
for packet in read_packets(sock): for packet in read_packets(sock):
@ -75,7 +75,7 @@ def run_live_capture(
if matched: if matched:
kind = session.pairs[-1][7] if session.pairs else "permanent" kind = session.pairs[-1][7] if session.pairs else "permanent"
label = POOL_META.get(kind, POOL_META["permanent"])["name"] label = POOL_META.get(kind, POOL_META["permanent"])["name"]
print(f"Captured {label} page {session.last_page_seen}") console.print_page_captured(label, session.last_page_seen)
finally: finally:
try: try:
sock.ioctl(socket.SIO_RCVALL, socket.RCVALL_OFF) sock.ioctl(socket.SIO_RCVALL, socket.RCVALL_OFF)
@ -86,14 +86,13 @@ def run_live_capture(
exports = [] exports = []
for kind in session.kinds_seen(): for kind in session.kinds_seen():
pairs = session.pairs_for_kind(kind) pairs = session.pairs_for_kind(kind)
best_run = session.best_run(kind) best_run, run_warnings = select_continuous_run_from_page_1(pairs)
rows = session.build_rows(kind) rows = session.build_rows(kind)
if kind == "arc_miracle_box": if kind == "arc_miracle_box":
_arc_run, gap_warnings = select_continuous_arc_run(pairs) warnings = run_warnings + arc_stability_warnings(rows)
warnings = gap_warnings + arc_stability_warnings(rows)
else: else:
rows, warnings = annotate_groups(rows, starts_from_page_1=bool(best_run and best_run[0][0] == 1)) rows, group_warnings = annotate_groups(rows)
warnings = page_gap_warnings(pairs, best_run) + warnings warnings = run_warnings + group_warnings
pages_seen = [p[0] for p in best_run] pages_seen = [p[0] for p in best_run]
csv_path, json_path = export_paths(kind) csv_path, json_path = export_paths(kind)
if write_debug_csv: if write_debug_csv:
@ -116,21 +115,40 @@ def run_live_capture(
} }
) )
console.print_results_header()
if not exports: if not exports:
print("No Monopoly history pages were captured.") console.print_problem("No history pages were captured.")
console.print_note("This tool must already be running when you press Start on the")
console.print_note("game's main menu. Log out to the main menu, enter the game")
console.print_note("again, then reopen the history screen.")
return {"exports": []} return {"exports": []}
if copy_clipboard and len(exports) == 1: all_warnings = []
payload = exports[0]["payload"] for item in exports:
copy_to_clipboard(payload) scan = item["export"]["scan"]
print("Export copied to clipboard.") console.print_export_summary(
elif len(exports) > 1: item["export"]["banner"]["name"],
print("Multiple banners captured; clipboard copy skipped.") scan["decoded_records"],
scan["exported_records"],
scan["skipped_records"],
)
for warning in scan["warnings"]:
console.print_warning(warning["code"], warning["reason"], warning.get("records"))
all_warnings.extend(scan["warnings"])
console.maybe_print_incomplete_hint(all_warnings)
print()
for item in exports: for item in exports:
if item["csv_path"] is not None: if item["csv_path"] is not None:
print(f"CSV written: {item['csv_path']}") console.print_note(f"CSV written: {item['csv_path']}")
print(f"Export written: {item['json_path']}") console.print_note(f"Export written: {item['json_path']}")
if copy_clipboard and len(exports) == 1:
copy_to_clipboard(exports[0]["payload"])
console.print_success("Export copied to clipboard - paste it straight into your tracker.")
elif len(exports) > 1:
console.print_note("Multiple banners captured; clipboard copy skipped so one export")
console.print_note("does not overwrite another.")
return {"exports": exports} return {"exports": exports}

View file

@ -11,7 +11,7 @@ from nte_history_exporter.decoder.arc import (
parse_arc_response, parse_arc_response,
select_continuous_arc_run, select_continuous_arc_run,
) )
from nte_history_exporter.decoder.boundary import longest_monotonic_page_run from nte_history_exporter.decoder.boundary import select_continuous_run_from_page_1
from nte_history_exporter.decoder.protocol import ( from nte_history_exporter.decoder.protocol import (
history_request_kind, history_request_kind,
is_history_request, is_history_request,
@ -147,7 +147,5 @@ class LiveHistorySession:
def best_run(self, kind: str | None = None) -> list[tuple]: def best_run(self, kind: str | None = None) -> list[tuple]:
pairs = self.pairs_for_kind(kind) if kind else self.pairs pairs = self.pairs_for_kind(kind) if kind else self.pairs
if kind == "arc_miracle_box": best_run, _warnings = select_continuous_run_from_page_1(pairs)
best_run, _warnings = select_continuous_arc_run(pairs)
return best_run return best_run
return longest_monotonic_page_run(pairs)

View file

@ -14,12 +14,13 @@ if str(SRC) not in sys.path:
from nte_history_exporter.decoder.boundary import annotate_groups, make_uid from nte_history_exporter.decoder.boundary import annotate_groups, make_uid
from nte_history_exporter.constants import LIMITED_CHARACTER_MARKER, MARKER from nte_history_exporter.constants import LIMITED_CHARACTER_MARKER, MARKER
from nte_history_exporter.decoder.boundary import page_gap_warnings from nte_history_exporter.decoder.boundary import select_continuous_run_from_page_1
from nte_history_exporter.decoder.protocol import decode_response_records, history_request_kind from nte_history_exporter.decoder.protocol import decode_response_records, history_request_kind
from nte_history_exporter.constants import POOL_META from nte_history_exporter.constants import POOL_META
from nte_history_exporter.mappings import ARC_META, CHARACTERS, ITEMS, REWARDS_BY_ID from nte_history_exporter.mappings import ARC_META, CHARACTERS, ITEMS, REWARDS_BY_ID
from nte_history_exporter.decoder.protocol import decode_reward_key, infer_reward_type from nte_history_exporter.decoder.protocol import decode_reward_key, infer_reward_type
from nte_history_exporter.decoder.arc import ( from nte_history_exporter.decoder.arc import (
arc_stability_warnings,
build_arc_rows_from_pairs, build_arc_rows_from_pairs,
decode_arc_key, decode_arc_key,
decode_arc_timestamp, decode_arc_timestamp,
@ -164,7 +165,7 @@ class BoundaryExportTests(unittest.TestCase):
def test_pages_1_to_5_exports_dice_complete_oldest_group(self): def test_pages_1_to_5_exports_dice_complete_oldest_group(self):
rows = load_reference_csv("monopoly_history_poc_13_pages_1_to_5_v4.csv") rows = load_reference_csv("monopoly_history_poc_13_pages_1_to_5_v4.csv")
annotated, warnings = annotate_groups(rows, starts_from_page_1=True) annotated, warnings = annotate_groups(rows)
exported = [row for row in annotated if row["export_record"] is True] exported = [row for row in annotated if row["export_record"] is True]
skipped = [row for row in annotated if row["export_record"] is False] skipped = [row for row in annotated if row["export_record"] is False]
@ -188,49 +189,75 @@ class BoundaryExportTests(unittest.TestCase):
"quantity": 1, "quantity": 1,
} }
def test_oldest_group_with_partial_dice_count_still_drops(self): def test_oldest_group_with_partial_dice_count_exports_with_warning(self):
rows = [self._synthetic_row(1, "aa", "dice") for _ in range(5)] rows = [self._synthetic_row(1, "aa", "dice") for _ in range(5)]
rows += [self._synthetic_row(2, "bb", "dice") for _ in range(5)] rows += [self._synthetic_row(2, "bb", "dice") for _ in range(5)]
rows += [self._synthetic_row(3, "bb", "dice") for _ in range(4)] rows += [self._synthetic_row(3, "bb", "dice") for _ in range(4)]
rows += [self._synthetic_row(3, "bb", "points_gift")] rows += [self._synthetic_row(3, "bb", "points_gift")]
annotated, warnings = annotate_groups(rows, starts_from_page_1=True) annotated, warnings = annotate_groups(rows)
exported = [row for row in annotated if row["export_record"] is True] exported = [row for row in annotated if row["export_record"] is True]
# Oldest group has 9 dice + 1 gift across a full final page: the gift does # Oldest group is an unfinished 10-pull on a full final page. Its captured
# not count toward pull-set sizing, so the group cannot be proven complete. # prefix is ordinal-stable, so it is exported with an informational warning
self.assertEqual(len(exported), 5) # rather than dropped.
self.assertEqual(len(exported), 15)
self.assertEqual(len(warnings), 1) self.assertEqual(len(warnings), 1)
self.assertEqual(warnings[0]["code"], "INCOMPLETE_TIMESTAMP_GROUP_EXPORTED")
self.assertEqual(warnings[0]["dice_records"], 9) self.assertEqual(warnings[0]["dice_records"], 9)
oldest = [row for row in annotated if row["timestamp_raw_hex"] == "bb"]
self.assertTrue(all(row["uid"] for row in oldest))
self.assertEqual([row["timestamp_group_ordinal"] for row in oldest], list(range(10)))
def test_incomplete_oldest_prefix_keeps_stable_uids(self):
full = [self._synthetic_row(1, "aa", "dice") for _ in range(5)]
full += [self._synthetic_row(2, "bb", "dice") for _ in range(3)]
full += [self._synthetic_row(3, "bb", "dice") for _ in range(2)]
truncated = [r for r in full if r["page"] in (1, 2)]
full_rows, _ = annotate_groups([dict(r) for r in full])
trunc_rows, _ = annotate_groups([dict(r) for r in truncated])
full_uids = [r["uid"] for r in full_rows if r["timestamp_raw_hex"] == "bb"][:3]
trunc_uids = [r["uid"] for r in trunc_rows if r["timestamp_raw_hex"] == "bb"]
# Capturing only the first 3 of a 5-record oldest group yields the same
# UIDs those rows have in the full capture.
self.assertEqual(len(trunc_uids), 3)
self.assertEqual(trunc_uids, full_uids)
def test_oldest_group_with_ten_dice_exports_on_full_final_page(self): def test_oldest_group_with_ten_dice_exports_on_full_final_page(self):
rows = [self._synthetic_row(1, "aa", "dice") for _ in range(5)] rows = [self._synthetic_row(1, "aa", "dice") for _ in range(5)]
rows += [self._synthetic_row(2, "bb", "dice") for _ in range(5)] rows += [self._synthetic_row(2, "bb", "dice") for _ in range(5)]
rows += [self._synthetic_row(3, "bb", "dice") for _ in range(5)] rows += [self._synthetic_row(3, "bb", "dice") for _ in range(5)]
annotated, warnings = annotate_groups(rows, starts_from_page_1=True) annotated, warnings = annotate_groups(rows)
exported = [row for row in annotated if row["export_record"] is True] exported = [row for row in annotated if row["export_record"] is True]
self.assertEqual(len(exported), 15) self.assertEqual(len(exported), 15)
self.assertEqual(len(warnings), 0) self.assertEqual(len(warnings), 0)
def test_newest_group_stays_dropped_mid_history_even_if_dice_complete(self): def test_run_selection_anchors_to_page_1_and_keeps_newest(self):
rows = [self._synthetic_row(3, "aa", "dice") for _ in range(5)] # Page 2's response was lost: captured pages 1, 3, 4, 5.
rows += [self._synthetic_row(4, "aa", "dice") for _ in range(5)] pairs = [(p, p * 2, 0, 0, 0, 0, b"", "permanent") for p in (1, 3, 4, 5)]
rows += [self._synthetic_row(5, "cc", "dice") for _ in range(4)] run, warnings = select_continuous_run_from_page_1(pairs)
annotated, warnings = annotate_groups(rows, starts_from_page_1=False) # The page-1 run (just page 1, the newest history) is kept; later pages are
exported = [row for row in annotated if row["export_record"] is True] # ignored with a gap warning, never silently discarding page 1.
self.assertEqual([p[0] for p in run], [1])
# Scan started mid-history: even 10 seen dice rolls cannot prove the newest
# group is whole, because newer same-timestamp rows would shift ordinals.
self.assertEqual(len(exported), 4)
self.assertEqual(len(warnings), 1) self.assertEqual(len(warnings), 1)
self.assertEqual(warnings[0]["dice_records"], 10) self.assertEqual(warnings[0]["code"], "PAGE_GAP_DETECTED")
self.assertEqual(warnings[0]["ignored_pages"], [3, 4, 5])
def test_run_selection_warns_when_page_1_missing(self):
pairs = [(p, p * 2, 0, 0, 0, 0, b"", "permanent") for p in (3, 4, 5)]
run, warnings = select_continuous_run_from_page_1(pairs)
self.assertEqual([p[0] for p in run], [3, 4, 5])
self.assertEqual(warnings[0]["code"], "DID_NOT_START_AT_PAGE_1")
def test_full_reference_scan_exports_all_rows(self): def test_full_reference_scan_exports_all_rows(self):
rows = load_reference_csv("monopoly_history_poc_10_all_44_pages_v4.csv") rows = load_reference_csv("monopoly_history_poc_10_all_44_pages_v4.csv")
annotated, warnings = annotate_groups(rows, starts_from_page_1=True) annotated, warnings = annotate_groups(rows)
exported = [row for row in annotated if row["export_record"] is True] exported = [row for row in annotated if row["export_record"] is True]
json_path = EXPORTS / "monopoly_history_export_10_all_44_pages_v4.json" json_path = EXPORTS / "monopoly_history_export_10_all_44_pages_v4.json"
@ -247,7 +274,7 @@ class BoundaryExportTests(unittest.TestCase):
def test_sanitized_export_omits_raw_packet_fields(self): def test_sanitized_export_omits_raw_packet_fields(self):
rows = load_reference_csv("monopoly_history_poc_13_pages_1_to_5_v4.csv") rows = load_reference_csv("monopoly_history_poc_13_pages_1_to_5_v4.csv")
annotated, warnings = annotate_groups(rows, starts_from_page_1=True) annotated, warnings = annotate_groups(rows)
export = build_export_json(annotated, warnings) export = build_export_json(annotated, warnings)
self.assertEqual(export["format"], "nte-history-export") self.assertEqual(export["format"], "nte-history-export")
@ -314,12 +341,10 @@ class BoundaryExportTests(unittest.TestCase):
self.assertEqual(make_uid(decoded, int(reference["timestamp_group_ordinal"])), "7d035ec098f856f81b403ea538810145") self.assertEqual(make_uid(decoded, int(reference["timestamp_group_ordinal"])), "7d035ec098f856f81b403ea538810145")
def test_page_gap_warning_reports_ignored_pages(self): def test_page_gap_warning_reports_ignored_pages(self):
pairs = [(1,), (2,), (3,), (5,)] pairs = [(p, p * 2, 0, 0, 0, 0, b"", "permanent") for p in (1, 2, 3, 5)]
best_run = [(1,), (2,), (3,)] run, warnings = select_continuous_run_from_page_1(pairs)
warnings = page_gap_warnings(pairs, best_run) self.assertEqual([p[0] for p in run], [1, 2, 3])
self.assertEqual(warnings[0]["code"], "PAGE_GAP_DETECTED") self.assertEqual(warnings[0]["code"], "PAGE_GAP_DETECTED")
self.assertEqual(warnings[0]["previous_page"], 3)
self.assertEqual(warnings[0]["next_page"], 5)
self.assertEqual(warnings[0]["ignored_pages"], [5]) self.assertEqual(warnings[0]["ignored_pages"], [5])
def test_arc_key_timestamp_and_uid_match_reference(self): def test_arc_key_timestamp_and_uid_match_reference(self):
@ -341,7 +366,7 @@ class BoundaryExportTests(unittest.TestCase):
self.assertEqual(decoded[0]["reward_type"], "arc") self.assertEqual(decoded[0]["reward_type"], "arc")
self.assertEqual(decoded[0]["reward_key_hex"], reference_rows[0]["arc_key_hex"]) self.assertEqual(decoded[0]["reward_key_hex"], reference_rows[0]["arc_key_hex"])
def test_arc_partial_timestamp_group_is_skipped(self): def test_arc_partial_timestamp_group_is_exported_with_warning(self):
rows = load_arc_csv("arc_pages_1_to_5_v2.csv") rows = load_arc_csv("arc_pages_1_to_5_v2.csv")
pairs = [] pairs = []
for page in range(1, 6): for page in range(1, 6):
@ -350,10 +375,34 @@ class BoundaryExportTests(unittest.TestCase):
pairs.append((page, page * 2, page, 1.0, page + 100, 1.1, response)) pairs.append((page, page * 2, page, 1.0, page + 100, 1.1, response))
decoded = build_arc_rows_from_pairs(pairs) decoded = build_arc_rows_from_pairs(pairs)
exported = [row for row in decoded if row["export_record"] is True] exported = [row for row in decoded if row["export_record"] is True]
skipped = [row for row in decoded if row["export_record"] is False] incomplete = [row for row in decoded if row["uid_status"] == "incomplete_stable"]
warnings = arc_stability_warnings(decoded)
# The oldest group is a 10-pull split by stopping at page 5 (5 of 10 rows).
# Its captured prefix is ordinal-stable, so it is exported, not dropped.
self.assertEqual(len(decoded), 25) self.assertEqual(len(decoded), 25)
self.assertEqual(len(exported), 20) self.assertEqual(len(exported), 25)
self.assertEqual(len(skipped), 5) self.assertEqual(len(incomplete), 5)
self.assertEqual(len(warnings), 1)
self.assertEqual(warnings[0]["code"], "INCOMPLETE_ARC_10_PULL_EXPORTED")
self.assertTrue(all(row["uid"] for row in incomplete))
def test_arc_incomplete_prefix_keeps_stable_uids(self):
rows = load_arc_csv("arc_pull_10_all_pages_v2.csv")
# Build a full 2-page (10-record) group, then a truncated 1-page version.
full_pairs, trunc_pairs = [], []
for page in (1, 2):
page_rows = [r for r in rows if int(r["page"]) == page]
response = bytes(0x4C) + b"".join(bytes.fromhex(r["record_hex"]) for r in page_rows)
full_pairs.append((page, page * 2, page, 1.0, page + 100, 1.1, response))
if page == 1:
trunc_pairs.append((page, page * 2, page, 1.0, page + 100, 1.1, response))
ts = rows[0]["timestamp_raw_hex"]
full = [r for r in build_arc_rows_from_pairs(full_pairs) if r["timestamp_raw_hex"] == ts]
trunc = [r for r in build_arc_rows_from_pairs(trunc_pairs) if r["timestamp_raw_hex"] == ts]
self.assertTrue(trunc)
self.assertEqual([r["uid"] for r in trunc], [r["uid"] for r in full[: len(trunc)]])
def test_arc_row_builder_accepts_live_pairs_with_kind(self): def test_arc_row_builder_accepts_live_pairs_with_kind(self):
reference_rows = load_arc_csv("arc_pull_10_all_pages_v2.csv")[:5] reference_rows = load_arc_csv("arc_pull_10_all_pages_v2.csv")[:5]
@ -419,7 +468,7 @@ class BoundaryExportTests(unittest.TestCase):
}, },
] ]
annotated, warnings = annotate_groups(rows, starts_from_page_1=True) annotated, warnings = annotate_groups(rows)
self.assertEqual([row["timestamp_group_ordinal"] for row in annotated[:4]], [0, 1, 2, 3]) self.assertEqual([row["timestamp_group_ordinal"] for row in annotated[:4]], [0, 1, 2, 3])
self.assertEqual({row["timestamp_group_size_seen"] for row in annotated[:4]}, {2}) self.assertEqual({row["timestamp_group_size_seen"] for row in annotated[:4]}, {2})