Search arXiv⌕ Search

arXiv subjects

Jonathan A. Karr Jr

Publications and source records attributed to Jonathan A. Karr Jr.

3 recordsLinked to original sources

Racing to the Starting Line: Measuring Opportunity and Improvement in Cross Country

Collegiate cross country programs set race schedules without large comparable datasets across courses and conditions. Using the comprehensive-era National Running Club Database (NRCD; 23,360 results; 7,056 athletes; 2023-2025), where course and weather coverage exceed 99%, we ask what public results can identify about seasonal improvement, measurement, and racing opportunity. Hierarchical models of Standardized times recover pooled within-season improvement of about -5.7 s/week for men and -5.2 for women, but reading those slopes as fitness requires residual meet difficulty not to track the calendar: crossed meet effects can erase the week slope, while repeating-venue effects recover a clearer men's estimate under a weaker assumption. After the opener, race-1-only forecasts of who will improve are near noise (held-out R^2 ~ 0). Weather standardization remains a useful convention (Converted Only overstates first-to-last gains by about 15-21 s relative to Standardized), and prospectively trained venue factors modestly help where prior course data exist. Successful programs race more and place higher at nationals, but that may not be why they win: the schedule-placement link is between programs and is entangled with roster size, not evidence that changing one team's race count raises nationals place.

cs.CY↗

Why AI Detection Fails for Academic Integrity

Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both as misconduct. In a controlled study of published English abstracts (four domains; 2013 to 2015 vs. 2023 to 2025), we quantify this policy failure under proxy human/AI labels at tau=0.50. Light "refine abstract only" edits, a proxy for guideline-compliant AI assistance, are flagged at 38 to 80%. Unmodified 2023 to 2025 originals are flagged at 9 to 15%, with non-STEM rates far above STEM (p<0.001); elevated scores track long-token and Academic Word List density, not authorship intent alone. After Undetectable AI humanization, evasion is near-total: fewer than 4% of AI-labeled rewrites remain flagged (post-humanization detection rate <4%; FNR >96%). Honest AI-editing results in a higher sanction risk than humanizer-assisted evasion. Therefore, detector scores should not serve as standalone misconduct evidence.

cs.LG↗

KEO: Knowledge Extraction on OMIn via Knowledge Graphs and RAG for Safety-Critical Aviation Maintenance

We present Knowledge Extraction on OMIn (KEO), a domain-specific knowledge extraction and reasoning framework with large language models (LLMs) in safety-critical contexts. Using the Operations and Maintenance Intelligence (OMIn) dataset, we construct a QA benchmark spanning global sensemaking and actionable maintenance tasks. KEO builds a structured Knowledge Graph (KG) and integrates it into a retrieval-augmented generation (RAG) pipeline, enabling more coherent, dataset-wide reasoning than traditional text-chunk RAG. We evaluate locally deployable LLMs (Gemma-3, Phi-4, Mistral-Nemo) and employ stronger models (GPT-4o, Llama-3.3) as judges. Experiments show that KEO markedly improves global sensemaking by revealing patterns and system-level insights, while text-chunk RAG remains effective for fine-grained procedural tasks requiring localized retrieval. These findings underscore the promise of KG-augmented LLMs for secure, domain-specific QA and their potential in high-stakes reasoning. The code is available at https://github.com/JonathanKarr33/keo.

cs.CL↗