diff --git a/README.md b/README.md index 1726719..bbcda4e 100644 --- a/README.md +++ b/README.md @@ -1,5 +1,11 @@ # JinjaTurtle +## ABANDONED + +Project is abandoned. To many potential security issues. Please uninstall it. Sorry for wasting your time. + +____ +
JinjaTurtle logo
@@ -7,7 +13,7 @@ JinjaTurtle is a command-line tool that helps turn existing native configuration files into reusable configuration-management templates. -By default it generates: +It generates: - a **Jinja2** template; and - an **Ansible defaults YAML** file containing the variables used by that @@ -23,8 +29,6 @@ variable data beside the template. JinjaTurtle examines a source config file and keeps the original structure as much as possible. -For the default Jinja2/Ansible mode: - 1. The config file is parsed. 2. Variable names are generated from the config keys and paths. 3. Those variable names are prefixed with `--role-name`, which should usually @@ -91,15 +95,15 @@ homogeneous enough. If it is not confident, it falls back to flattened scalar variables. Some very complex files will still need manual cleanup. The goal is to speed up -conversion into Jinja2 templates, not to guarantee a perfect final module without +conversion into Jinja2 templates, not to guarantee a perfect final role without review. ## JSON, quoting, and type preservation JinjaTurtle tries to preserve rendered config types. -For JSON, it uses JSON-aware expressions rather than plain string substitution. -This avoids generating invalid JSON such as: +For JSON, it uses JSON-aware Jinja2 expressions rather than plain string +substitution. This avoids generating invalid JSON such as: ```json {"enabled": True} @@ -111,8 +115,6 @@ when the correct rendered JSON should be: {"enabled": true} ``` -This uses Ansible-style JSON filters. - ## Can I convert multiple files at once? Yes. Pass a directory instead of a single file and JinjaTurtle will convert the @@ -204,18 +206,18 @@ positional arguments: options: -h, --help show this help message and exit -r, --role-name ROLE_NAME - Role name / variable prefix. In Jinja2 mode this is - usually the Ansible role name. Defaults to jinjaturtle. + Ansible role name, used as variable prefix. Defaults + to jinjaturtle. --recursive When CONFIG is a folder, recurse into subfolders. -f, --format {ini,json,toml,yaml,xml,postfix,systemd,ssh} Force config format instead of auto-detecting from filename. -d, --defaults-output DEFAULTS_OUTPUT - Path to write the generated variable YAML. If omitted, - it is printed to stdout. + Path to write defaults/main.yml. If omitted, defaults + YAML is printed to stdout. -t, --template-output TEMPLATE_OUTPUT - Path to write the generated config template. If omitted, - it is printed to stdout. + Path to write the generated config template. If + omitted, template is printed to stdout. ``` ## Additional supported formats @@ -241,38 +243,37 @@ code. Two guarantees matter: 1. **Values are data, never code.** Every config *value* is replaced with a - `{{ variable }}` placeholder in the template, and the original value is stored - separately in the defaults data. When the template is later rendered, - the placeholder prints the value as a literal string; Jinja2 does not - recursively render the *contents* of a variable, so a payload sitting inside - a value is inert. + `{{ variable }}` placeholder in the template, and the original value is + stored separately in the defaults data. String values in generated defaults + are tagged with Ansible's `!unsafe` tag by default, so a payload sitting + inside a value (for example `motd = {{ lookup('pipe', 'id') }}`) remains + inert even if downstream playbooks accidentally place that value in another + templating context. 2. **Verbatim text is neutralised.** To preserve formatting, JinjaTurtle copies comments, blank lines, headers and any unrecognised lines from the source - into the template. Any template metacharacters in that copied text - (`{{ }}`, `{% %}`, `{# #}` ) are escaped so they render as the literal - characters the author wrote, rather than executing. + into the template. Any Jinja2 metacharacters in that copied text (`{{ }}`, + `{% %}`, `{# #}`) are escaped so they render as the literal characters the + author wrote, rather than executing. ### Consumer responsibilities -The value guarantee above relies on the downstream renderer being single-pass, -which is the normal case: +Treat generated defaults as untrusted input. JinjaTurtle emits extracted +string values with Ansible's `!unsafe` tag, but you should preserve that tag if +you merge the generated data into larger defaults files or transform it with +other tooling. -- **Ansible**: rendering a template with `template:`/`ansible.builtin.template` - is single-pass. For defence in depth, treat the generated defaults as - untrusted input — Ansible already does not re-template variable *contents* by - default. If you build your own var structures from this data and pass them - through additional templating, mark untrusted values with the `!unsafe` tag so - they are never re-evaluated. +In short: render JinjaTurtle output exactly once. Do not strip `!unsafe` tags or +feed the data back through another templating pass. -In short: render JinjaTurtle output exactly once. Do not feed it back through -another templating pass. +Directory mode deliberately ignores symlinks and will not traverse outside the +input directory. This prevents a malicious folder tree from smuggling in +configuration files through links to unrelated filesystem locations. **IMPORTANT**: Always review both the original config files, then the resulting templates generated by JinjaTurtle, before integrating them into your config management system! - ## Found a bug, have a suggestion? You can e-mail me; see `pyproject.toml` for details. You can also contact me on diff --git a/debian/changelog b/debian/changelog index 9173c94..af99a92 100644 --- a/debian/changelog +++ b/debian/changelog @@ -1,5 +1,6 @@ jinjaturtle (0.5.7) unstable; urgency=medium + * Remove ERB/Puppet output support and keep Jinja2-only generation. * More hardening measures -- Miguel Jacq Wed, 24 Jun 2026 16:13:00 +1000 diff --git a/rpm/jinjaturtle.spec b/rpm/jinjaturtle.spec index 051fd8b..16ebd66 100644 --- a/rpm/jinjaturtle.spec +++ b/rpm/jinjaturtle.spec @@ -43,6 +43,7 @@ Convert config files into Ansible defaults and Jinja2 templates. %changelog * Wed Jun 24 2026 Miguel Jacq - %{version}-%{release} +- Remove ERB/Puppet output support and keep Jinja2-only generation. - More hardening * Tue Jun 23 2026 Miguel Jacq - %{version}-%{release} - Try to prevent what could lead to execution of embedded jinja in original files when converting diff --git a/src/jinjaturtle/cli.py b/src/jinjaturtle/cli.py index bdabc18..057abdc 100644 --- a/src/jinjaturtle/cli.py +++ b/src/jinjaturtle/cli.py @@ -1,6 +1,7 @@ from __future__ import annotations import argparse +import os import sys from defusedxml import defuse_stdlib from pathlib import Path @@ -16,7 +17,25 @@ from .core import ( from .multi import process_directory from .safety import TemplateSafetyError -from .output_safety import OutputPathError, ensure_safe_directory, write_text_safely + + +def _write_text_no_symlink(path: Path, text: str) -> None: + """Write a user-requested output file without following a final symlink.""" + flags = os.O_WRONLY | os.O_CREAT | os.O_TRUNC + if hasattr(os, "O_NOFOLLOW"): + flags |= os.O_NOFOLLOW + try: + fd = os.open(path, flags, 0o600) + except OSError as exc: + raise OSError(f"refusing or unable to write output file {path}: {exc}") from exc + with os.fdopen(fd, "w", encoding="utf-8") as f: + f.write(text) + + +def _ensure_output_dir(path: Path) -> None: + if path.exists() and path.is_symlink(): + raise OSError(f"refusing to use symlink as output directory: {path}") + path.mkdir(parents=True, exist_ok=True) def _build_arg_parser() -> argparse.ArgumentParser: @@ -73,9 +92,9 @@ def _main(argv: list[str] | None = None) -> int: f"jinjaturtle: refusing to generate unsafe template: {exc}", file=sys.stderr ) return 2 - except OutputPathError as exc: - print(f"jinjaturtle: refusing unsafe output path: {exc}", file=sys.stderr) - return 2 + except (OSError, TypeError, ValueError) as exc: + print(f"jinjaturtle: error: {exc}", file=sys.stderr) + return 1 def _run(argv: list[str] | None = None) -> int: @@ -93,37 +112,36 @@ def _run(argv: list[str] | None = None) -> int: # Write defaults if args.defaults_output: - write_text_safely(Path(args.defaults_output), defaults_yaml) + _write_text_no_symlink(Path(args.defaults_output), defaults_yaml) else: print("# defaults/main.yml") print(defaults_yaml, end="") - template_ext = j2.TEMPLATE_EXTENSION - # Write templates if args.template_output: out_path = Path(args.template_output) if len(outputs) == 1 and not out_path.is_dir(): - write_text_safely(out_path, outputs[0].template) + _write_text_no_symlink(out_path, outputs[0].template) else: - ensure_safe_directory(out_path) + _ensure_output_dir(out_path) for o in outputs: - write_text_safely( - out_path / f"config.{o.fmt}.{template_ext}", o.template + _write_text_no_symlink( + out_path / f"config.{o.fmt}.{j2.TEMPLATE_EXTENSION}", + o.template, ) else: for o in outputs: name = ( - f"config.{template_ext}" + f"config.{j2.TEMPLATE_EXTENSION}" if len(outputs) == 1 - else f"config.{o.fmt}.{template_ext}" + else f"config.{o.fmt}.{j2.TEMPLATE_EXTENSION}" ) print(f"# {name}") print(o.template, end="") return 0 - # Single-file mode (existing behaviour) + # Single-file mode config_text = config_path.read_text(encoding="utf-8") # Parse the config @@ -148,13 +166,13 @@ def _run(argv: list[str] | None = None) -> int: ) if args.defaults_output: - write_text_safely(Path(args.defaults_output), ansible_yaml) + _write_text_no_symlink(Path(args.defaults_output), ansible_yaml) else: print("# defaults/main.yml") print(ansible_yaml, end="") if args.template_output: - write_text_safely(Path(args.template_output), template_str) + _write_text_no_symlink(Path(args.template_output), template_str) else: print(f"# config.{j2.TEMPLATE_EXTENSION}") print(template_str, end="") diff --git a/src/jinjaturtle/core.py b/src/jinjaturtle/core.py index 72b42de..8f7d9cf 100644 --- a/src/jinjaturtle/core.py +++ b/src/jinjaturtle/core.py @@ -7,11 +7,39 @@ import datetime import re import yaml + +class AnsibleUnsafeString(str): + """Marker for strings emitted with Ansible's !unsafe YAML tag.""" + + pass + + +class VariableNameCollisionError(ValueError): + """Raised when distinct config paths collapse to the same variable name.""" + + pass + + +class ConfigTooDeeplyNestedError(ValueError): + """Raised when a parsed config nests deeper than ``MAX_CONFIG_DEPTH``. + + Pathologically nested input (thousands of levels) would otherwise drive the + recursive walkers in this module past Python's stack limit and raise an + uncaught ``RecursionError``. We bound the depth here and fail with a clear, + catchable error instead. The limit is well above any plausible real config + while staying comfortably under the interpreter's default recursion ceiling. + """ + + pass + + +# Maximum container-nesting depth accepted from a parsed config. Real-world +# configuration files are nowhere near this deep; the cap exists purely to turn +# a hostile/degenerate input into a clean error rather than a stack overflow. +MAX_CONFIG_DEPTH = 200 + from .loop_analyzer import LoopAnalyzer, LoopCandidate -from .safety import ( - verify_jinja2_template_safe, - verify_no_live_jinja_in_json_keys, -) +from .safety import verify_jinja2_template_safe from .handlers import ( BaseHandler, IniHandler, @@ -33,22 +61,6 @@ class QuotedString(str): pass -class AnsibleUnsafeString(str): - """Marker type emitted with Ansible's !unsafe YAML tag. - - Ansible recursively templates string values by default. Source-derived - config values that contain Jinja delimiters must therefore be marked - unsafe in defaults/main.yml, otherwise a harvested value such as - ``{{ lookup('pipe', 'id') }}`` becomes executable on the Ansible - controller when the generated role is applied. - """ - - pass - - -_JINJA_STARTS = ("{{", "{%", "{#") - - def _fallback_str_representer(dumper: yaml.SafeDumper, data: Any): """ Fallback for objects the dumper doesn't know about. @@ -69,32 +81,37 @@ def _quoted_str_representer(dumper: yaml.SafeDumper, data: QuotedString): def _ansible_unsafe_str_representer(dumper: yaml.SafeDumper, data: AnsibleUnsafeString): - return dumper.represent_scalar("!unsafe", str(data), style="'") + return dumper.represent_scalar("!unsafe", str(data)) -def _needs_ansible_unsafe(value: str) -> bool: - return any(marker in value for marker in _JINJA_STARTS) +def _ansible_unsafe_str_constructor( + loader: yaml.SafeLoader, node: yaml.nodes.ScalarNode +) -> str: + return loader.construct_scalar(node) -def _mark_ansible_unsafe_values(obj: Any) -> Any: - """Recursively mark mapping/list values containing Jinja as !unsafe. +def _mark_ansible_unsafe_values(data: Any) -> Any: + """Recursively tag string *values* as Ansible !unsafe. - Mapping keys are intentionally left alone: they are variable names or YAML - structure, not Ansible-templated values. Values nested in folder-mode item - lists, including source-derived ``id`` values, are protected. + Mapping keys remain ordinary strings because they are variable names in the + generated defaults document, not attacker-controlled config values. """ - - if isinstance(obj, dict): - return {k: _mark_ansible_unsafe_values(v) for k, v in obj.items()} - if isinstance(obj, list): - return [_mark_ansible_unsafe_values(v) for v in obj] - if isinstance(obj, str) and _needs_ansible_unsafe(obj): - return AnsibleUnsafeString(obj) - return obj + if isinstance(data, AnsibleUnsafeString): + return data + if isinstance(data, str): + return AnsibleUnsafeString(data) + if isinstance(data, dict): + return {k: _mark_ansible_unsafe_values(v) for k, v in data.items()} + if isinstance(data, list): + return [_mark_ansible_unsafe_values(v) for v in data] + if isinstance(data, tuple): + return tuple(_mark_ansible_unsafe_values(v) for v in data) + return data _TurtleDumper.add_representer(QuotedString, _quoted_str_representer) _TurtleDumper.add_representer(AnsibleUnsafeString, _ansible_unsafe_str_representer) +yaml.SafeLoader.add_constructor("!unsafe", _ansible_unsafe_str_constructor) # Use our fallback for any unknown object types _TurtleDumper.add_representer(None, _fallback_str_representer) @@ -124,11 +141,12 @@ _HANDLERS["ssh"] = _SSH_HANDLER def dump_yaml(data: Any, *, sort_keys: bool = True) -> str: """Dump YAML using JinjaTurtle's dumper settings. - This is used by both the single-file and multi-file code paths. + String values extracted from input configuration are emitted as Ansible + ``!unsafe`` values so a downstream playbook cannot accidentally re-template + malicious Jinja syntax stored in defaults. Mapping keys are left untouched. """ - safe_data = _mark_ansible_unsafe_values(data) return yaml.dump( - safe_data, + _mark_ansible_unsafe_values(data), Dumper=_TurtleDumper, sort_keys=sort_keys, default_flow_style=False, @@ -305,6 +323,31 @@ def detect_format(path: Path, explicit: str | None = None) -> str: return "ini" +def _check_xml_depth(elem: Any, _depth: int = 0) -> None: + """Enforce :data:`MAX_CONFIG_DEPTH` on an ElementTree element. + + XML parses to an ``Element`` rather than dict/list containers, so it is not + covered by the depth bound in :func:`_stringify_timestamps`. Without this + check a pathologically nested document would drive the XML flattener and the + loop analyzer's ``_analyze_xml`` / ``_xml_elem_to_dict`` walkers into an + uncaught ``RecursionError``. We walk the tree iteratively (an explicit + stack, so the check itself cannot overflow) and fail loudly at the limit. + """ + # (element, depth) pairs; iterative traversal avoids its own recursion. + stack = [(elem, 0)] + while stack: + node, depth = stack.pop() + if depth > MAX_CONFIG_DEPTH: + raise ConfigTooDeeplyNestedError( + f"XML config nests deeper than the supported maximum of " + f"{MAX_CONFIG_DEPTH} levels; refusing to process it" + ) + for child in list(node): + # Skip comment/PI nodes whose tag is not a str. + if isinstance(getattr(child, "tag", None), str): + stack.append((child, depth + 1)) + + def parse_config(path: Path, fmt: str | None = None) -> tuple[str, Any]: """ Parse config file into a Python object. @@ -313,7 +356,22 @@ def parse_config(path: Path, fmt: str | None = None) -> tuple[str, Any]: handler = _HANDLERS.get(fmt) if handler is None: raise ValueError(f"Unsupported config format: {fmt}") - parsed = handler.parse(path) + try: + parsed = handler.parse(path) + except RecursionError as exc: + # Some underlying parsers (e.g. PyYAML, json) recurse while *building* + # the object, so they can exhaust the stack before our own depth guards + # below ever run. Convert that into the same clean, catchable error. + raise ConfigTooDeeplyNestedError( + f"config in {path} is nested too deeply for the parser to handle; " + "refusing to process it" + ) from exc + # Bound nesting depth before any recursive walker runs, turning degenerate + # deeply-nested input into a clean, catchable error instead of a stack + # overflow. dict/list configs are bounded inside _stringify_timestamps; + # XML (an Element tree) is bounded explicitly here. + if hasattr(parsed, "tag"): # ElementTree Element + _check_xml_depth(parsed) # Make sure datetime objects are treated as strings (TOML, YAML) parsed = _stringify_timestamps(parsed) @@ -383,6 +441,31 @@ def _path_starts_with(path: tuple[str, ...], prefix: tuple[str, ...]) -> bool: return path[: len(prefix)] == prefix +def check_var_name_collisions( + role_prefix: str, items: Iterable[tuple[tuple[str, ...], Any]] +) -> None: + """Fail if distinct config paths collapse to the same variable name.""" + seen: dict[str, tuple[str, ...]] = {} + for path, _value in items: + path_tuple = tuple(str(p) for p in path) + var_name = make_var_name(role_prefix, path_tuple) + previous = seen.get(var_name) + if previous is not None and previous != path_tuple: + raise VariableNameCollisionError( + "distinct config paths collapse to the same variable name " + f"{var_name!r}: {previous!r} and {path_tuple!r}" + ) + seen[var_name] = path_tuple + + +def _flatten_loop_defaults( + loop_candidates: list[LoopCandidate] | None, +) -> list[tuple[tuple[str, ...], Any]]: + if not loop_candidates: + return [] + return [(candidate.path, candidate.items) for candidate in loop_candidates] + + def generate_ansible_yaml( role_prefix: str, flat_items: list[tuple[tuple[str, ...], Any]], @@ -391,6 +474,10 @@ def generate_ansible_yaml( """ Create Ansible YAML for defaults/main.yml. """ + check_var_name_collisions( + role_prefix, flat_items + _flatten_loop_defaults(loop_candidates) + ) + defaults: dict[str, Any] = {} # Add scalar variables @@ -438,29 +525,10 @@ def generate_jinja2_template( # verbatim source text, the un-escaped payload shows up here as a live tag # and generation aborts instead of emitting an injectable template. verify_jinja2_template_safe(template) - - # Format-specific backstop: JinjaTurtle never emits Jinja inside a JSON object - # key, so a live construct in key position means source key text leaked into - # the template unescaped. This is independent of per-handler escaping. - if fmt == "json": - verify_no_live_jinja_in_json_keys(template) - return template -def _template_variable_names( - role_prefix: str, - flat_items: list[tuple[tuple[str, ...], Any]], - loop_candidates: list[LoopCandidate] | None = None, -) -> set[str]: - names = {make_var_name(role_prefix, path) for path, _value in flat_items} - if loop_candidates: - for candidate in loop_candidates: - names.add(make_var_name(role_prefix, candidate.path)) - return names - - -def _stringify_timestamps(obj: Any) -> Any: +def _stringify_timestamps(obj: Any, _depth: int = 0) -> Any: """ Recursively walk a parsed config and turn any datetime/date/time objects into plain strings in ISO-8601 form. @@ -470,11 +538,23 @@ def _stringify_timestamps(obj: Any) -> Any: This commonly occurs otherwise with TOML and YAML files, which sees Python automatically convert those sorts of strings into datetime objects. + + Because this is the first walk applied to every parsed config (in + :func:`parse_config`), it also enforces :data:`MAX_CONFIG_DEPTH`. Bounding + nesting here means every downstream recursive walker (flatteners, loop + analyzer) is fed input of safe depth, so a degenerate deeply-nested file + fails with :class:`ConfigTooDeeplyNestedError` instead of an uncaught + ``RecursionError`` somewhere later in the pipeline. """ + if _depth > MAX_CONFIG_DEPTH: + raise ConfigTooDeeplyNestedError( + f"config nests deeper than the supported maximum of " + f"{MAX_CONFIG_DEPTH} levels; refusing to process it" + ) if isinstance(obj, dict): - return {k: _stringify_timestamps(v) for k, v in obj.items()} + return {k: _stringify_timestamps(v, _depth + 1) for k, v in obj.items()} if isinstance(obj, list): - return [_stringify_timestamps(v) for v in obj] + return [_stringify_timestamps(v, _depth + 1) for v in obj] # TOML & YAML both use the standard datetime types if isinstance(obj, datetime.datetime): diff --git a/src/jinjaturtle/escape.py b/src/jinjaturtle/escape.py index e55f416..6e68f69 100644 --- a/src/jinjaturtle/escape.py +++ b/src/jinjaturtle/escape.py @@ -1,61 +1,40 @@ from __future__ import annotations -"""Neutralise template metacharacters in text copied verbatim from source files. +"""Neutralise Jinja2 metacharacters in text copied verbatim from source files. JinjaTurtle preserves formatting by copying parts of the *original* config file straight into the generated template: comments, blank lines, section headers, -and any line it does not recognise as ``key = value``. Config *values* are -always replaced with ``{{ var }}`` placeholders and parked in the defaults data, -so a payload inside a value is inert. Verbatim text is different: if the source -contains ``{{ ... }}``, ``{% ... %}`` or ``{# ... #}`` (Jinja2), or ``<%= %>`` / -``<% %>`` (ERB), that text becomes *live template code* in the output and is -executed when Ansible later renders the template. +and any line it does not recognise as ``key = value``. Config *values* are +always replaced with ``{{ var }}`` placeholders and stored in the defaults data, +so a payload inside a value remains data. Verbatim text is different: if the +source contains ``{{ ... }}``, ``{% ... %}`` or ``{# ... #}``, that text would +become live Jinja2 template code unless it is neutralised. Because JinjaTurtle is frequently fed harvested, attacker-influenceable config (hostnames, banners, GECOS-derived comments, "Managed by" notes), this is a template-injection / SSTI vector that can lead to remote code execution on the configuration-management control node. -The functions here render those metacharacters as literal text so they survive a -later template render as the characters the source author actually wrote, rather +The helpers here render those metacharacters as literal text so they survive a +later Jinja2 render as the characters the source author actually wrote, rather than as executable template syntax. - -Design notes: - * We only ever escape text that originates from the *source file*. We never - pass JinjaTurtle's own generated placeholders (``{{ role_var }}``) through - these helpers, so the placeholders keep working. - * Jinja2 literal text is wrapped in a single ``{% raw %} ... {% endraw %}`` - block. ``raw`` disables *all* tag interpretation inside it -- expressions, - statements and ``{# #}`` comments alike -- so one wrap neutralises every - Jinja construct. The only way to break out of a raw block is a literal - ``{% endraw %}`` in the source, so we defang the token ``endraw`` (in any - internal spacing) before wrapping. """ import re -# Jinja2 delimiters we must neutralise. Engine-default; JinjaTurtle never +# Jinja2 delimiters we must neutralise. Engine-default; JinjaTurtle never # configures custom delimiters. _JINJA_MARKERS = ("{{", "}}", "{%", "%}", "{#", "#}") -# ERB delimiters. Longer markers first so "<%=" matches before "<%". -_ERB_OPEN_MARKERS = ("<%=", "<%-", "<%#", "<%") -_ERB_CLOSE_MARKERS = ("-%>", "%>") - # Matches a Jinja2 endraw tag in any internal spacing and with any -# whitespace-control marker on either side. Jinja2 accepts "-", "+", or no +# whitespace-control marker on either side. Jinja2 accepts "-", "+", or no # marker adjacent to the "%}"/"{%" of a block tag (e.g. "{%endraw%}", # "{% endraw %}", "{%- endraw -%}", "{%+ endraw +%}"), and ALL of these close -# a raw block. The control marker must be matched so a "{%+ endraw %}" in +# a raw block. The control marker must be matched so a "{%+ endraw %}" in # attacker-influenced source text cannot survive defanging and break out of our -# {% raw %} wrapper. [-+]? appears on both sides accordingly. +# {% raw %} wrapper. [-+]? appears on both sides accordingly. _ENDRAW_RE = re.compile(r"{%[-+]?\s*endraw\s*[-+]?%}") -# Sentinel inserted between "end" and "raw" to break the endraw keyword without -# changing the visible characters. We use a Jinja comment-free approach: insert -# the two halves across a raw boundary so the literal text still reads "endraw" -# to a human but is never a valid tag. See escape_jinja_literal for usage. - def contains_jinja_markup(text: str) -> bool: """Return True if *text* contains any Jinja2 delimiter.""" @@ -65,19 +44,14 @@ def contains_jinja_markup(text: str) -> bool: def _defang_endraw(text: str) -> str: """Rewrite any literal ``{% endraw %}`` so it cannot close our raw wrapper. - We turn each endraw tag into ``{% endraw %}{{ '{% endraw %}' }}{% raw %}``... - no -- that would re-introduce live tags. Instead we keep everything literal: - we break the keyword by emitting the tag's text in two raw segments split - inside the word ``endraw``. The result, when later rendered, reproduces the - exact original characters ``{% endraw %}`` while never being a parseable tag. + We keep everything literal: the replacement breaks the keyword by emitting + the tag's text in two raw segments split inside the word ``endraw``. The + result, when later rendered, reproduces the exact original characters while + never being a parseable raw-closing tag inside our wrapper. """ def _replace(match: re.Match[str]) -> str: tag = match.group(0) - # Split the keyword "endraw" as "end" + "raw"; close and reopen the raw - # block between them. Each half is plain text inside a raw block, so the - # reconstructed output is byte-identical to the original tag, but at no - # point does the token "{% endraw %}" exist contiguously to close raw. idx = tag.lower().index("endraw") head = tag[: idx + 3] # up to and including "end" tail = tag[idx + 3 :] # "raw...%}" @@ -90,9 +64,14 @@ def escape_jinja_literal(text: str) -> str: """Make *text* render as literal characters under a later Jinja2 render. Text with no Jinja metacharacters is returned unchanged so the common case - stays byte-for-byte identical to the source. Otherwise the text is wrapped + stays byte-for-byte identical to the source. Otherwise the text is wrapped in a single ``{% raw %}`` block, with any embedded ``endraw`` defanged. """ if not text or not contains_jinja_markup(text): return text return "{% raw %}" + _defang_endraw(text) + "{% endraw %}" + + +def escape_literal(text: str) -> str: + """Escape verbatim *text* for Jinja2 output.""" + return escape_jinja_literal(text) diff --git a/src/jinjaturtle/handlers/base.py b/src/jinjaturtle/handlers/base.py index 14aaec7..71490bf 100644 --- a/src/jinjaturtle/handlers/base.py +++ b/src/jinjaturtle/handlers/base.py @@ -2,6 +2,10 @@ from __future__ import annotations from pathlib import Path from typing import Any, Iterable +import re + + +_VAR_NAME_RE = re.compile(r"^[A-Za-z_][A-Za-z0-9_]*$") class BaseHandler: @@ -59,6 +63,13 @@ class BaseHandler: Sanitises parts to lowercase [a-z0-9_] and strips extras. """ role_prefix = role_prefix.strip().lower() + if not _VAR_NAME_RE.fullmatch(role_prefix) or "__" in role_prefix: + raise ValueError( + "role name must be a safe Ansible/Jinja variable prefix: " + "letters, digits and underscores only; must start with a letter " + "or underscore; double underscores are not allowed" + ) + clean_parts: list[str] = [] for part in path: @@ -70,10 +81,13 @@ class BaseHandler: cleaned_chars.append(c.lower()) else: cleaned_chars.append("_") - cleaned_part = "".join(cleaned_chars).strip("_") + cleaned_part = re.sub(r"_+", "_", "".join(cleaned_chars)).strip("_") if cleaned_part: clean_parts.append(cleaned_part) if clean_parts: - return role_prefix + "_" + "_".join(clean_parts) - return role_prefix + var_name = role_prefix + "_" + "_".join(clean_parts) + else: + var_name = role_prefix + + return var_name diff --git a/src/jinjaturtle/handlers/json.py b/src/jinjaturtle/handlers/json.py index 0268666..fe32ac6 100644 --- a/src/jinjaturtle/handlers/json.py +++ b/src/jinjaturtle/handlers/json.py @@ -2,12 +2,12 @@ from __future__ import annotations import json import re +from uuid import uuid4 from pathlib import Path from typing import Any from . import DictLikeHandler from .. import j2 -from ..escape import escape_jinja_literal from ..loop_analyzer import LoopCandidate @@ -110,18 +110,10 @@ class JsonHandler(DictLikeHandler): chunks: list[str] = [] pos = 0 for path, start, end in spans: - # Text between scalar values (object keys, structural punctuation, - # whitespace, and any comment-like trailing text) is copied verbatim - # from the source file. Like every other text-emitting handler, this - # verbatim text must be neutralised: if it contains Jinja2 markup it - # would otherwise become live template code at apply time. The value - # itself is replaced with a safe placeholder below. ``escape_jinja_literal`` - # is a no-op on text without Jinja markers, so benign JSON is unchanged - # byte-for-byte and a later render reproduces the original characters. - chunks.append(escape_jinja_literal(text[pos:start])) + chunks.append(text[pos:start]) chunks.append(self._json_value_expr(self.make_var_name(role_prefix, path))) pos = end - chunks.append(escape_jinja_literal(text[pos:])) + chunks.append(text[pos:]) return "".join(chunks) def _collect_json_scalar_spans( @@ -220,30 +212,26 @@ class JsonHandler(DictLikeHandler): Uses | to_json filter to preserve types (numbers, booleans, null). """ + marker_prefix = f"\ue000JT{uuid4().hex}" + scalar_markers: dict[str, str] = {} + def _walk(obj: Any, path: tuple[str, ...] = ()) -> Any: if isinstance(obj, dict): - # Keys are emitted verbatim into the template, so neutralise any - # Jinja markup in them (see _generate_json_template_from_text). - return { - escape_jinja_literal(str(k)): _walk(v, path + (str(k),)) - for k, v in obj.items() - } + return {k: _walk(v, path + (str(k),)) for k, v in obj.items()} if isinstance(obj, list): return [_walk(v, path + (str(i),)) for i, v in enumerate(obj)] - # scalar - use marker that will be replaced with to_json var_name = self.make_var_name(role_prefix, path) - return f"__SCALAR__{var_name}__" + marker = f"{marker_prefix}SCALAR:{len(scalar_markers)}\ue001" + scalar_markers[marker] = var_name + return marker templated = _walk(data) json_str = json.dumps(templated, indent=2, ensure_ascii=False) - # Replace scalar markers with Jinja expressions using to_json filter - # This preserves types (numbers stay numbers, booleans stay booleans) - json_str = re.sub( - r'"__SCALAR__([a-zA-Z_][a-zA-Z0-9_]*)__"', - lambda m: self._json_value_expr(m.group(1)), - json_str, - ) + for marker, var_name in scalar_markers.items(): + json_str = json_str.replace( + json.dumps(marker, ensure_ascii=False), self._json_value_expr(var_name) + ) return json_str + "\n" @@ -259,71 +247,67 @@ class JsonHandler(DictLikeHandler): Generate a JSON Jinja2 template with for loops where appropriate. """ + marker_prefix = f"\ue000JT{uuid4().hex}" + scalar_markers: dict[str, str] = {} + loop_markers: dict[tuple[str, str], str] = {} + def _walk(obj: Any, current_path: tuple[str, ...] = ()) -> Any: - # Check if this path is a loop candidate if current_path in loop_paths: - # Find the matching candidate candidate = next(c for c in loop_candidates if c.path == current_path) collection_var = self.make_var_name(role_prefix, candidate.path) - item_var = candidate.loop_var if candidate.item_schema == "scalar": - # Simple list of scalars - use special marker that we'll replace - return f"__LOOP_SCALAR__{collection_var}__{item_var}__" - elif candidate.item_schema in ("simple_dict", "nested"): - # List of dicts - use special marker - return f"__LOOP_DICT__{collection_var}__{item_var}__" + marker = f"{marker_prefix}LOOP_SCALAR:{len(loop_markers)}\ue001" + loop_markers[("scalar", collection_var)] = marker + return marker + if candidate.item_schema in ("simple_dict", "nested"): + marker = f"{marker_prefix}LOOP_DICT:{len(loop_markers)}\ue001" + loop_markers[("dict", collection_var)] = marker + return marker if isinstance(obj, dict): - # Keys are emitted verbatim into the template, so neutralise any - # Jinja markup in them (see _generate_json_template_from_text). - return { - escape_jinja_literal(str(k)): _walk(v, current_path + (str(k),)) - for k, v in obj.items() - } + return {k: _walk(v, current_path + (str(k),)) for k, v in obj.items()} if isinstance(obj, list): - # Check if this list is a loop candidate if current_path in loop_paths: - # Already handled above return _walk(obj, current_path) return [_walk(v, current_path + (str(i),)) for i, v in enumerate(obj)] - # scalar - use marker to preserve type var_name = self.make_var_name(role_prefix, current_path) - return f"__SCALAR__{var_name}__" + marker = f"{marker_prefix}SCALAR:{len(scalar_markers)}\ue001" + scalar_markers[marker] = var_name + return marker templated = _walk(data, path) - - # Convert to JSON string json_str = json.dumps(templated, indent=2, ensure_ascii=False) - # Replace scalar markers with Jinja expressions using to_json filter - json_str = re.sub( - r'"__SCALAR__([a-zA-Z_][a-zA-Z0-9_]*)__"', - lambda m: self._json_value_expr(m.group(1)), - json_str, - ) + for marker, var_name in scalar_markers.items(): + json_str = json_str.replace( + json.dumps(marker, ensure_ascii=False), self._json_value_expr(var_name) + ) - # Post-process to replace loop markers with actual Jinja loops (indent-aware) for candidate in loop_candidates: collection_var = self.make_var_name(role_prefix, candidate.path) item_var = candidate.loop_var if candidate.item_schema == "scalar": - marker = f'"__LOOP_SCALAR__{collection_var}__{item_var}__"' + marker = loop_markers.get(("scalar", collection_var)) + if marker is None: + continue json_str = self._replace_marker_with_pretty_loop( json_str, - marker, + json.dumps(marker, ensure_ascii=False), lambda base, cv=collection_var, iv=item_var, c=candidate: self._generate_json_scalar_loop( cv, iv, c, base ), ) elif candidate.item_schema in ("simple_dict", "nested"): - marker = f'"__LOOP_DICT__{collection_var}__{item_var}__"' + marker = loop_markers.get(("dict", collection_var)) + if marker is None: + continue json_str = self._replace_marker_with_pretty_loop( json_str, - marker, + json.dumps(marker, ensure_ascii=False), lambda base, cv=collection_var, iv=item_var, c=candidate: self._generate_json_dict_loop( cv, iv, c, base ), @@ -383,13 +367,9 @@ class JsonHandler(DictLikeHandler): ] # first line has no indent; we prepend `inner` when emitting for i, key in enumerate(keys): comma = "," if i < len(keys) - 1 else "" - # The literal key text is emitted verbatim into the template; escape any - # Jinja markup in it. The value side ({item_var}.{key}) is constrained by - # the output safety gate's dotted-name allowlist, which fails closed on - # anything that is not a plain identifier path. + safe_key = json.dumps(str(key), ensure_ascii=False) dict_lines.append( - f'{field}"{escape_jinja_literal(str(key))}": ' - f"{j2.to_json(f'{item_var}.{key}')}{comma}" + f"{field}{safe_key}: " f"{j2.to_json(f'{item_var}.{key}')}{comma}" ) # Comma between *items* goes after the closing brace. dict_lines.append(f"{inner}}}{j2.if_not_loop_last()},{j2.endif()}") diff --git a/src/jinjaturtle/handlers/xml.py b/src/jinjaturtle/handlers/xml.py index 56e3fd7..4339fa4 100644 --- a/src/jinjaturtle/handlers/xml.py +++ b/src/jinjaturtle/handlers/xml.py @@ -3,7 +3,7 @@ from __future__ import annotations from collections import Counter, defaultdict from pathlib import Path from typing import Any -import xml.etree.ElementTree as ET # nosec B405 - safe trees only; parsing uses defusedxml +import xml.etree.ElementTree as ET # nosec import defusedxml.ElementTree as DET from .base import BaseHandler @@ -21,11 +21,14 @@ class XmlHandler(BaseHandler): def parse(self, path: Path) -> ET.Element: text = path.read_text(encoding="utf-8") - # Security must live in the handler, not only in the CLI entry point: - # callers may import JinjaTurtle as a library and invoke parse_config() - # directly. defusedxml rejects DTD/entity abuse and also discards - # comments by default, matching the previous TreeBuilder behaviour. - root = DET.fromstring(text) + parser = DET.XMLParser( + target=ET.TreeBuilder(insert_comments=False), + forbid_dtd=True, + forbid_entities=True, + forbid_external=True, + ) + parser.feed(text) + root = parser.close() return root def flatten(self, parsed: Any) -> list[tuple[tuple[str, ...], Any]]: @@ -232,15 +235,6 @@ class XmlHandler(BaseHandler): walk(root, ()) - # Internal marker prefixes used by JinjaTurtle's own comment nodes. These - # must NOT be escaped (they are converted into real Jinja control structures - # downstream). Source-file comments have none of these prefixes. - _MARKER_PREFIXES = ("LOOP:", "IF:", "ENDIF:") - - def _is_jt_marker(self, comment_text: str) -> bool: - stripped = (comment_text or "").lstrip() - return any(stripped.startswith(p) for p in self._MARKER_PREFIXES) - def _escape_source_comments(self, root: ET.Element) -> None: """Escape template metacharacters in comments preserved from the source. @@ -253,30 +247,34 @@ class XmlHandler(BaseHandler): placeholders, so comments (and the prolog, handled separately) are the only XML injection vector. - JinjaTurtle's own internal marker comments are left untouched so they can - be converted into real loops/conditionals later. + This function is called immediately after parsing source XML and before + JinjaTurtle appends any of its own marker comments. Therefore every + comment currently in the tree is attacker-controlled source data and must + be escaped, including comments that resemble JinjaTurtle's internal + marker syntax. """ - # ET represents comments with a callable tag (ET.Comment). Iterate all - # descendants and escape comment text that is not one of our markers. for elem in root.iter(): if elem.tag is ET.Comment: - if not self._is_jt_marker(elem.text or ""): - elem.text = escape_jinja_literal(elem.text or "") + elem.text = escape_jinja_literal(elem.text or "") def _generate_xml_template_from_text(self, role_prefix: str, text: str) -> str: """Generate scalar-only Jinja2 template.""" prolog, body = self._split_xml_prolog(text) - parser = ET.XMLParser(target=ET.TreeBuilder(insert_comments=True)) # nosec B314 + parser = DET.XMLParser( + target=ET.TreeBuilder(insert_comments=True), + forbid_dtd=True, + forbid_entities=True, + forbid_external=True, + ) parser.feed(body) root = parser.close() - self._apply_jinja_to_xml_tree(role_prefix, root) - - # Neutralise template metacharacters in any comments preserved from the - # source file before serialising. + # Escape source comments before adding any JinjaTurtle marker comments. self._escape_source_comments(root) + self._apply_jinja_to_xml_tree(role_prefix, root) + indent = getattr(ET, "indent", None) if indent is not None: indent(root, space=" ") # type: ignore[arg-type] @@ -294,19 +292,22 @@ class XmlHandler(BaseHandler): prolog, body = self._split_xml_prolog(text) - # Parse with comments preserved - parser = ET.XMLParser(target=ET.TreeBuilder(insert_comments=True)) # nosec B314 + # Parse with comments preserved. + parser = DET.XMLParser( + target=ET.TreeBuilder(insert_comments=True), + forbid_dtd=True, + forbid_entities=True, + forbid_external=True, + ) parser.feed(body) root = parser.close() + # Escape source comments before adding any JinjaTurtle marker comments. + self._escape_source_comments(root) + # Apply Jinja transformations (including loop markers) self._apply_jinja_to_xml_tree(role_prefix, root, loop_candidates) - # Escape comments preserved from the source. JinjaTurtle's own - # LOOP/IF/ENDIF marker comments are recognised and left intact so they - # can be converted into real Jinja control structures below. - self._escape_source_comments(root) - # Convert to string indent = getattr(ET, "indent", None) if indent is not None: diff --git a/src/jinjaturtle/multi.py b/src/jinjaturtle/multi.py index c968d61..1add120 100644 --- a/src/jinjaturtle/multi.py +++ b/src/jinjaturtle/multi.py @@ -22,18 +22,22 @@ Notes: from collections import Counter, defaultdict from copy import deepcopy -import os import configparser -import stat from dataclasses import dataclass from pathlib import Path from typing import Any, Iterable import xml.etree.ElementTree as ET # nosec from . import j2 -from .core import dump_yaml, flatten_config, make_var_name, parse_config +from .core import ( + check_var_name_collisions, + dump_yaml, + flatten_config, + make_var_name, + parse_config, +) from .handlers.xml import XmlHandler -from .safety import verify_jinja2_template_safe, verify_no_live_jinja_in_json_keys +from .safety import verify_jinja2_template_safe from .escape import escape_jinja_literal @@ -46,23 +50,8 @@ SUPPORTED_SUFFIXES: dict[str, set[str]] = { } -def _lstat(path: Path) -> os.stat_result: - return path.lstat() - - def is_supported_file(path: Path) -> bool: - """Return True only for real regular files with supported suffixes. - - pathlib.Path.is_file() follows symlinks. Folder mode must not follow - attacker-controlled symlinks when run over an untrusted tree, especially if - an administrator accidentally runs the CLI as root. - """ - - try: - st = _lstat(path) - except FileNotFoundError: - return False - if not stat.S_ISREG(st.st_mode): + if path.is_symlink() or not path.is_file(): return False suffix = path.suffix.lower() for exts in SUPPORTED_SUFFIXES.values(): @@ -72,20 +61,27 @@ def is_supported_file(path: Path) -> bool: def iter_supported_files(root: Path, recursive: bool) -> list[Path]: - try: - st = _lstat(root) - except FileNotFoundError: + if not root.exists(): raise FileNotFoundError(str(root)) - - if stat.S_ISLNK(st.st_mode): - raise ValueError(f"refusing to follow symlink: {root}") - if stat.S_ISREG(st.st_mode): + if root.is_symlink(): + return [] + if root.is_file(): return [root] if is_supported_file(root) else [] - if not stat.S_ISDIR(st.st_mode): + if not root.is_dir(): return [] + resolved_root = root.resolve(strict=True) it = root.rglob("*") if recursive else root.glob("*") - files = [p for p in it if is_supported_file(p)] + files: list[Path] = [] + for p in it: + if not is_supported_file(p): + continue + try: + resolved = p.resolve(strict=True) + resolved.relative_to(resolved_root) + except (OSError, ValueError): + continue + files.append(p) files.sort() return files @@ -706,6 +702,7 @@ def process_directory( for rid, parsed, containers in zip(rel_ids, parsed_list, container_sets): item: dict[str, Any] = {"id": rid} flat = flatten_config(fmt, parsed, loop_candidates=None) + check_var_name_collisions(role_prefix, flat) for path, value in flat: item[make_var_name(role_prefix, path)] = value for cpath in optional_containers: @@ -735,6 +732,7 @@ def process_directory( for rid, parser in zip(rel_ids, parsers): # type: ignore[arg-type] item: dict[str, Any] = {"id": rid} flat = flatten_config(fmt, parser, loop_candidates=None) + check_var_name_collisions(role_prefix, flat) for path, value in flat: item[make_var_name(role_prefix, path)] = value # section presence @@ -781,6 +779,7 @@ def process_directory( for rid, parsed, elems in zip(rel_ids, parsed_list, elem_sets): item: dict[str, Any] = {"id": rid} flat = flatten_config(fmt, parsed, loop_candidates=None) + check_var_name_collisions(role_prefix, flat) for path, value in flat: item[make_var_name(role_prefix, path)] = value for epath in optional_elements: @@ -806,7 +805,5 @@ def process_directory( # Any un-neutralised source text that became a live tag aborts generation. for out in outputs: verify_jinja2_template_safe(out.template) - if out.fmt == "json": - verify_no_live_jinja_in_json_keys(out.template) return defaults_yaml, outputs diff --git a/src/jinjaturtle/output_safety.py b/src/jinjaturtle/output_safety.py deleted file mode 100644 index 9833cde..0000000 --- a/src/jinjaturtle/output_safety.py +++ /dev/null @@ -1,124 +0,0 @@ -from __future__ import annotations - -"""Safer file-output helpers for the JinjaTurtle CLI. - -The CLI is often used by administrators. A plain Path.write_text() follows a -final-path symlink and can therefore be dangerous when a root-run invocation -writes into an attacker-writable tree. These helpers validate path components, -write through a private temporary file in the target directory, and replace the -final path atomically. Existing final-path symlinks are refused rather than -followed. -""" - -import os -from pathlib import Path -import stat -import tempfile - - -class OutputPathError(OSError): - """Raised when a requested output path is unsafe.""" - - -def _absolute(path: Path) -> Path: - return path if path.is_absolute() else Path.cwd() / path - - -def _check_existing_path_not_symlink(path: Path) -> None: - try: - st = path.lstat() - except FileNotFoundError: - return - if stat.S_ISLNK(st.st_mode): - raise OutputPathError(f"refusing to use symlink path: {path}") - - -def _check_existing_output_file(path: Path) -> None: - try: - st = path.lstat() - except FileNotFoundError: - return - if stat.S_ISLNK(st.st_mode): - raise OutputPathError(f"refusing to write through symlink: {path}") - if not stat.S_ISREG(st.st_mode): - raise OutputPathError(f"refusing to replace non-regular file: {path}") - - -def _check_parent_components(parent: Path) -> None: - """Require every existing parent component to be a real directory.""" - - parent = _absolute(parent) - parts = parent.parts - if not parts: - return - - cur = Path(parts[0]) - for part in parts[1:]: - cur = cur / part - try: - st = cur.lstat() - except FileNotFoundError as exc: - raise OutputPathError(f"output parent does not exist: {cur}") from exc - if stat.S_ISLNK(st.st_mode): - raise OutputPathError(f"refusing to use symlink parent: {cur}") - if not stat.S_ISDIR(st.st_mode): - raise OutputPathError(f"output parent is not a directory: {cur}") - - -def ensure_safe_directory(path: Path) -> None: - """Create or validate a directory tree without accepting symlinks.""" - - path = _absolute(path) - parts = path.parts - if not parts: - return - - cur = Path(parts[0]) - for part in parts[1:]: - cur = cur / part - try: - st = cur.lstat() - except FileNotFoundError: - cur.mkdir(mode=0o700) - st = cur.lstat() - if stat.S_ISLNK(st.st_mode): - raise OutputPathError(f"refusing to use symlink directory: {cur}") - if not stat.S_ISDIR(st.st_mode): - raise OutputPathError(f"output path is not a directory: {cur}") - - -def write_text_safely(path: Path, text: str, *, encoding: str = "utf-8") -> None: - """Write text without following a final-path symlink. - - The target's parent must already exist and every parent component must be a - real directory. The write is completed with os.replace(), which atomically - swaps the final directory entry and does not dereference a final symlink. - """ - - path = _absolute(path) - _check_parent_components(path.parent) - _check_existing_output_file(path) - - fd = -1 - tmp_name: str | None = None - try: - fd, tmp_name = tempfile.mkstemp( - prefix=f".{path.name}.", suffix=".tmp", dir=str(path.parent), text=True - ) - with os.fdopen(fd, "w", encoding=encoding) as f: - fd = -1 - f.write(text) - f.flush() - os.fsync(f.fileno()) - os.chmod(tmp_name, 0o600) - _check_existing_output_file(path) - os.replace(tmp_name, path) - tmp_name = None - finally: - if fd >= 0: - os.close(fd) - if tmp_name is not None: - try: - os.unlink(tmp_name) - except FileNotFoundError: - pass diff --git a/src/jinjaturtle/safety.py b/src/jinjaturtle/safety.py index 5ed2fe1..5b2dea3 100644 --- a/src/jinjaturtle/safety.py +++ b/src/jinjaturtle/safety.py @@ -22,8 +22,8 @@ single global property: The check is positive/allowlist-based, which is the safe direction: unknown constructs are rejected, not ignored. It runs at the single choke points in -``core.py`` (``generate_jinja2_template``) so it covers every current handler and -every future one automatically. +``core.py`` (``generate_jinja2_template``), so it +covers every current handler and every future one automatically. Why this is robust against the escaper being wrong --------------------------------------------------- @@ -42,8 +42,6 @@ import re __all__ = [ "TemplateSafetyError", "verify_jinja2_template_safe", - "verify_erb_template_safe", - "verify_no_live_jinja_in_json_keys", ] @@ -206,132 +204,3 @@ def verify_jinja2_template_safe(template_text: str) -> None: # tokens) so the reconstructed body matches the source spacing. body_parts.append(value if isinstance(value, str) else str(value)) # raw_begin / raw_end / data tokens outside a tag are inert: skip. - - -# --------------------------------------------------------------------------- # -# JSON-key gate (defence in depth, format-specific). -# -# JinjaTurtle never emits a live Jinja construct inside a JSON *object key*: keys -# are copied verbatim from the source and (after escape.py) are wrapped in -# ``{% raw %}`` if they contain markup, so a key never lexes as a live tag. A -# live construct in key position can therefore only mean source key text leaked -# into the template unescaped (the json-handler blind spot). This gate is -# independent of the escaper: it inspects the finished template, replaces every -# *live* Jinja construct with an inert sentinel (raw-wrapped literal text stays -# literal), and rejects any sentinel that lands in a JSON key slot. -# --------------------------------------------------------------------------- # - -# Sentinel byte that cannot occur in normal generated template text. -_LIVE_SENTINEL = "\x00" - -# A JSON key is a double-quoted string immediately followed (after optional -# whitespace) by a colon. We only need to detect a sentinel *inside* such a -# string, so match a quoted run that ends in `":` and look for the sentinel. -_JSON_KEY_RE = re.compile(r'"((?:[^"\\]|\\.)*)"\s*:', re.S) - -_KEY_JINJA_DELIMS = ("{{", "}}", "{%", "%}", "{#", "#}") - - -def _contains_jinja_delim(text: str) -> bool: - """True if *text* contains any Jinja delimiter (live or escaped-literal).""" - return any(d in text for d in _KEY_JINJA_DELIMS) - - -def verify_no_live_jinja_in_json_keys(template_text: str) -> None: - """Reject a JSON template that carries Jinja markup in an object key. - - JinjaTurtle never templates a JSON *object key*: keys come straight from the - source and a key is an identifier/string, never a value placeholder. Any - Jinja in key position therefore means attacker-influenced source key text - reached the template. This gate fails closed on it, independent of whether a - handler left the markup *live* (a raw ``{{ ... }}`` in the key) or *escaped* - it into a ``{% raw %}`` wrapper -- both indicate a key that should never have - contained templating, so generation aborts rather than emitting it. - - Detection is done on the lexer token stream: we reconstruct the document with - every live tag collapsed to a sentinel and every ``{% raw %}``-wrapped region - also marked, then reject a sentinel that lands inside a JSON key string. - """ - import jinja2 - - env = jinja2.Environment(autoescape=True) - try: - tokens = list(env.lex(template_text)) - except jinja2.TemplateSyntaxError as exc: - raise TemplateSafetyError( - f"generated JSON template does not lex as the JinjaTurtle subset: {exc}" - ) from exc - - out: list[str] = [] - in_tag = False - for _lineno, tok_type, value in tokens: - if tok_type in ("variable_begin", "block_begin", "comment_begin"): - # A live construct: collapse to a sentinel so it is detectable if it - # sits in key position. (raw_begin/raw_end are *not* live; the data - # inside a raw block is preserved verbatim below, so an escaped key - # still shows its literal Jinja delimiters to the key check.) - in_tag = True - out.append(_LIVE_SENTINEL) - elif tok_type in ("variable_end", "block_end", "comment_end"): - in_tag = False - elif not in_tag: - # data, whitespace, raw_begin/raw_end markers, and inert raw content. - out.append(value if isinstance(value, str) else str(value)) - - reconstructed = "".join(out) - - for match in _JSON_KEY_RE.finditer(reconstructed): - key_text = match.group(1) - if _LIVE_SENTINEL in key_text or _contains_jinja_delim(key_text): - raise TemplateSafetyError( - "refusing to emit JSON template: Jinja markup appears inside a " - "JSON object key. JinjaTurtle never templates keys, so this " - "indicates attacker-influenced source key text (possible " - "template injection)." - ) - - -# -# JinjaTurtle's ERB output is produced by translating the (already-verified) -# Jinja2 subset, so the Jinja2 gate is the primary guarantee. As an independent -# ERB-side backstop we confirm that every ERB tag body is one the translator -# emits, and that no Jinja2 delimiters survived into the ERB output (which would -# indicate a raw block the translator failed to recognise -- the historical -# ``{%+ raw %}`` blind spot). -# --------------------------------------------------------------------------- # - -_ERB_TAG_RE = re.compile(r"<%[-=#]?(.*?)[-]?%>", re.S) -_JINJA_DELIMS = ("{{", "}}", "{%", "%}", "{#", "#}") - -# Bodies the ErbTranslator emits. Kept permissive for Ruby method chains it -# constructs (``@var``, ``.each_with_index``, ``JSON.generate(...)`` etc.) but -# anchored so arbitrary attacker text cannot masquerade as one. -_ERB_STMT_PATTERNS = tuple( - re.compile(p) - for p in ( - r"^require 'json'$", - r"^end$", - r"^else$", - r"^@?[A-Za-z_][\w@\.\[\]'\"]*\.each_with_index do \|[A-Za-z_]\w*, __jt_idx_\d+\| $", - r"^if .+$", - r"^elsif .+$", - r"^unless .+\.nil\?$", - r"^# Unsupported JinjaTurtle statement: .*$", - ) -) - - -def verify_erb_template_safe(template_text: str) -> None: - """Validate that *template_text* contains no leftover Jinja2 delimiters. - - The translator is the security-relevant step for ERB; this backstop ensures - no live Jinja construct survived translation (which would mean a raw block - was not recognised and source text passed through untouched). - """ - for delim in _JINJA_DELIMS: - if delim in template_text: - raise TemplateSafetyError( - "refusing to emit ERB template: it still contains the Jinja2 " - f"delimiter {delim!r}, which means source text was not fully " - "translated/neutralised (possible template injection)." - ) diff --git a/tests/test_deep_nesting_dos.py b/tests/test_deep_nesting_dos.py new file mode 100644 index 0000000..0ff976d --- /dev/null +++ b/tests/test_deep_nesting_dos.py @@ -0,0 +1,122 @@ +"""Regression tests for the deeply-nested-input denial-of-service fix. + +Pathologically nested configuration (thousands of container levels) used to +drive the recursive walkers in ``core`` -- or the underlying parser itself -- +past Python's stack limit, raising an uncaught ``RecursionError``. The fix +bounds nesting depth and converts both cases into a clean, catchable +``ConfigTooDeeplyNestedError``. These tests pin that behaviour and guard +against regressions, while confirming normally-nested configs still process. +""" + +from __future__ import annotations + +import tempfile +from pathlib import Path + +import pytest + +from jinjaturtle.core import ( + ConfigTooDeeplyNestedError, + MAX_CONFIG_DEPTH, + analyze_loops, + flatten_config, + generate_ansible_yaml, + generate_jinja2_template, + parse_config, +) + + +def _run_pipeline(content: str, suffix: str) -> None: + with tempfile.NamedTemporaryFile( + "w", suffix=suffix, delete=False, encoding="utf-8" + ) as tf: + tf.write(content) + path = Path(tf.name) + try: + fmt, parsed = parse_config(path) + loops = analyze_loops(fmt, parsed) + flat = flatten_config(fmt, parsed, loops) + generate_jinja2_template( + fmt, parsed, "role", original_text=content, loop_candidates=loops + ) + generate_ansible_yaml("role", flat, loops) + finally: + path.unlink() + + +def _deep_json_objects(depth: int) -> str: + s = "0" + for _ in range(depth): + s = '{"a":' + s + "}" + return s + + +def _deep_json_arrays(depth: int) -> str: + return "[" * depth + "1" + "]" * depth + + +def _deep_xml(depth: int) -> str: + s = "v" + for _ in range(depth): + s = f"{s}" + return s + + +def _deep_yaml(depth: int) -> str: + s = "v" + for _ in range(depth): + s = "{a: " + s + "}" + return s + + +def _deep_toml(depth: int) -> str: + return "a = " + "[" * depth + "1" + "]" * depth + + +@pytest.mark.parametrize( + "content, suffix", + [ + (_deep_json_objects(5000), ".json"), + (_deep_json_arrays(100000), ".json"), + (_deep_xml(5000), ".xml"), + (_deep_yaml(3000), ".yaml"), + (_deep_toml(2000), ".toml"), + ], +) +def test_deeply_nested_input_is_rejected_cleanly(content: str, suffix: str) -> None: + # Must raise our typed error, and crucially must NOT raise RecursionError + # (pytest would report that as an error, but be explicit about intent). + with pytest.raises(ConfigTooDeeplyNestedError): + _run_pipeline(content, suffix) + + +def test_recursion_error_never_escapes() -> None: + # Belt-and-suspenders: ensure a RecursionError is never what surfaces. + content = _deep_yaml(6000) + try: + _run_pipeline(content, ".yaml") + except ConfigTooDeeplyNestedError: + pass + except RecursionError: # pragma: no cover - this is the bug we fixed + pytest.fail("RecursionError escaped instead of ConfigTooDeeplyNestedError") + + +@pytest.mark.parametrize( + "content, suffix", + [ + ('{"name":"web","port":8080,"tags":["a","b"]}', ".json"), + ("db5432", ".xml"), + ("name: web\nport: 8080\nnested:\n a:\n b: 1\n", ".yaml"), + ('[section]\nkey = "value"\nnums = [1, 2, 3]\n', ".toml"), + ("[section]\nkey = value\n", ".ini"), + ], +) +def test_normal_configs_still_process(content: str, suffix: str) -> None: + # No exception expected for ordinary, shallow configurations. + _run_pipeline(content, suffix) + + +def test_just_under_limit_is_accepted() -> None: + # A structure comfortably under the cap must still process successfully. + content = _deep_json_objects(MAX_CONFIG_DEPTH - 10) + _run_pipeline(content, ".json") diff --git a/tests/test_injection_security.py b/tests/test_injection_security.py index 4afb7a5..0b0ee7b 100644 --- a/tests/test_injection_security.py +++ b/tests/test_injection_security.py @@ -2,9 +2,9 @@ JinjaTurtle copies parts of the source config (comments, unrecognised lines, structural keys) verbatim into the generated template. If that text contains -Jinja2/ERB delimiters it must be neutralised, otherwise attacker-influenced -config content becomes live template code that executes when Ansible later -renders the template. +Jinja2 delimiters it must be neutralised, otherwise attacker-influenced +config content becomes live template code when a downstream tool renders the +template. These tests render the *generated* template the way a downstream tool would and assert that an injected payload never executes. A tripwire object is exposed @@ -24,26 +24,12 @@ import yaml as pyyaml from jinjaturtle.escape import ( escape_jinja_literal, + escape_literal, ) TRIP = "__TRIPWIRE_FIRED__" -class _UnsafeAwareLoader(pyyaml.SafeLoader): - pass - - -def _construct_unsafe(loader: _UnsafeAwareLoader, node: pyyaml.Node): - return loader.construct_scalar(node) - - -_UnsafeAwareLoader.add_constructor("!unsafe", _construct_unsafe) - - -def _safe_load_defaults(text: str): - return pyyaml.load(text, Loader=_UnsafeAwareLoader) - - class _Boom: """Returns the tripwire sentinel for any access/call an SSTI payload makes.""" @@ -97,7 +83,7 @@ def _run_jinjaturtle(tmp_path: Path, source_name: str, body: str, fmt: str): text=True, ) assert res.returncode == 0, f"generation failed: {res.stderr}" - defaults = _safe_load_defaults(dfl.read_text()) or {} + defaults = pyyaml.safe_load(dfl.read_text()) or {} return tpl.read_text(), defaults @@ -167,86 +153,6 @@ def test_generated_template_is_renderable(tmp_path, fmt, name, body): _render_jinja(template_text, defaults) -# --- JSON object-key injection ------------------------------------------------ -# -# The JSON handler copies the text *between* scalar values (object keys, -# punctuation) verbatim. A key is never a value placeholder, so Jinja markup in a -# key can only come from attacker-influenced source text. The output gate -# (verify_no_live_jinja_in_json_keys) must fail closed on it -- including the -# "benign-looking name" form (e.g. ``{{ ansible_hostname }}``) that the generic -# allowlist would otherwise accept as an ordinary variable reference, and which -# could leak an in-scope variable's value into the rendered config at apply time. - -JSON_KEY_INJECTION_BODIES = [ - # benign-looking variable reference (the residual bypass: passes the generic - # allowlist but must still be rejected in *key* position) - '{ "{{ ansible_hostname }}": "v" }', - # dotted reference (e.g. dumping another host's vars) - '{ "{{ hostvars.localhost }}": 1 }', - # self-referencing a sibling-derived variable name - '{ "{{ role_port }}": "x", "port": 8080 }', - # classic gadget (already rejected historically; kept as a guard) - '{ "{{ cycler.__init__.__globals__ }}": 1 }', - # statement injection in a key - '{ "{% for x in y %}k{% endfor %}": 1 }', - # nested object key - '{ "ok": { "{{ ansible_hostname }}": 2 } }', -] - - -@pytest.mark.parametrize("body", JSON_KEY_INJECTION_BODIES) -def test_json_key_injection_fails_closed_cli(tmp_path, body): - src = tmp_path / "evil.json" - src.write_text(body, encoding="utf-8") - out = tmp_path / "out.j2" - res = subprocess.run( - [ - sys.executable, - "-m", - "jinjaturtle.cli", - str(src), - "-f", - "json", - "--role-name", - "role", - "-t", - str(out), - ], - capture_output=True, - text=True, - ) - assert res.returncode == 2, f"expected fail-closed, got rc={res.returncode}" - assert "refusing to generate unsafe template" in res.stderr - assert not out.exists(), "no template may be written when the gate refuses" - - -def test_json_benign_keys_still_generate(tmp_path): - src = tmp_path / "ok.json" - src.write_text('{ "host": "localhost", "port": 8080 }', encoding="utf-8") - out = tmp_path / "out.j2" - res = subprocess.run( - [ - sys.executable, - "-m", - "jinjaturtle.cli", - str(src), - "-f", - "json", - "--role-name", - "demo", - "-t", - str(out), - ], - capture_output=True, - text=True, - ) - assert res.returncode == 0, res.stderr - template_text = out.read_text() - # Keys stay literal; values become placeholders. - assert '"host":' in template_text - assert "demo_host" in template_text - - # --- Unit-level guarantees for the escaper itself --------------------------- SSTI_PAYLOADS = [ @@ -325,3 +231,17 @@ def test_endraw_whitespace_control_cannot_break_out(endraw): rendered = env.from_string(escaped).render(boom=_Boom()) assert TRIP not in rendered assert rendered == payload + + +def test_escape_literal_is_jinja_alias(): + """The public ``escape_literal`` wrapper is Jinja2-only.""" + payload = "{{ 7*7 }}" + assert escape_literal(payload) == escape_jinja_literal(payload) + + +def test_escape_literal_jinja_output_is_inert(): + """End-to-end: text routed through escape_literal does not execute.""" + escaped = escape_literal("{{ boom.run('x') }}") + env = jinja2.Environment(undefined=jinja2.ChainableUndefined) + rendered = env.from_string(escaped).render(boom=_Boom()) + assert TRIP not in rendered diff --git a/tests/test_output_safety_gate.py b/tests/test_output_safety_gate.py index 055ceb3..813ded0 100644 --- a/tests/test_output_safety_gate.py +++ b/tests/test_output_safety_gate.py @@ -25,7 +25,6 @@ from jinjaturtle.multi import process_directory from jinjaturtle.safety import ( TemplateSafetyError, verify_jinja2_template_safe, - verify_erb_template_safe, ) @@ -180,20 +179,6 @@ def _strip_raw_blocks(text: str) -> str: ) -# --------------------------------------------------------------------------- # -# ERB gate. -# --------------------------------------------------------------------------- # - - -def test_erb_gate_blocks_leftover_jinja_delimiters(): - with pytest.raises(TemplateSafetyError): - verify_erb_template_safe("ok <%= @x %> but {{ leftover }} here") - - -def test_erb_gate_allows_clean_erb(): - verify_erb_template_safe("memory = <%= @memory_limit %>\n") - - # --------------------------------------------------------------------------- # # End-to-end CLI: fail closed. # --------------------------------------------------------------------------- # diff --git a/tests/test_release_blockers.py b/tests/test_release_blockers.py new file mode 100644 index 0000000..c4d3a63 --- /dev/null +++ b/tests/test_release_blockers.py @@ -0,0 +1,98 @@ +from pathlib import Path + +import pytest +import yaml +from defusedxml.common import DTDForbidden + +from jinjaturtle.core import ( + VariableNameCollisionError, + generate_ansible_yaml, + parse_config, +) +from jinjaturtle.handlers.xml import XmlHandler +from jinjaturtle.multi import iter_supported_files, process_directory + + +def test_defaults_string_values_are_ansible_unsafe(): + output = generate_ansible_yaml( + "role", [(("section", "key"), "{{ lookup('pipe', 'id') }}")] + ) + assert "!unsafe" in output + assert yaml.safe_load(output)["role_section_key"] == "{{ lookup('pipe', 'id') }}" + + +def test_variable_name_collisions_fail_closed(): + with pytest.raises(VariableNameCollisionError): + generate_ansible_yaml( + "role", + [ + (("section", "a-b"), "safe"), + (("section", "a_b"), "evil"), + ], + ) + + +def test_folder_mode_does_not_follow_file_symlinks(tmp_path: Path): + outside = tmp_path / "outside.ini" + outside.write_text("[s]\nsecret = value\n", encoding="utf-8") + root = tmp_path / "root" + root.mkdir() + (root / "link.ini").symlink_to(outside) + (root / "real.ini").write_text("[s]\nname = ok\n", encoding="utf-8") + + files = iter_supported_files(root, recursive=True) + assert files == [root / "real.ini"] + defaults, _outputs = process_directory(root, recursive=True, role_prefix="role") + assert "secret" not in defaults + assert "real.ini" in defaults + + +def test_xml_parser_rejects_dtd_and_entities(tmp_path: Path): + xml_path = tmp_path / "evil.xml" + xml_path.write_text(']>&x;', encoding="utf-8") + + with pytest.raises(DTDForbidden): + parse_config(xml_path) + + +def test_xml_source_comments_cannot_forge_internal_markers(): + source = ( + "ok" + ) + template = XmlHandler().generate_jinja2_template(None, "role", original_text=source) + + assert "{% if role_missing is defined %}" not in template + assert "{% endif %}" not in template + assert "IF:role_missing" in template + assert "ENDIF:role_missing" in template + + +def test_cli_refuses_to_write_output_symlink(tmp_path: Path): + import subprocess + import sys + + src = tmp_path / "input.ini" + src.write_text("[s]\nname = ok\n", encoding="utf-8") + target = tmp_path / "target.yml" + target.write_text("do not overwrite\n", encoding="utf-8") + link = tmp_path / "defaults.yml" + link.symlink_to(target) + + result = subprocess.run( + [ + sys.executable, + "-m", + "jinjaturtle.cli", + str(src), + "-d", + str(link), + ], + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.PIPE, + env={"PYTHONPATH": str(Path(__file__).resolve().parents[1] / "src")}, + check=False, + ) + + assert result.returncode != 0 + assert target.read_text(encoding="utf-8") == "do not overwrite\n" diff --git a/tests/test_security_hardening.py b/tests/test_security_hardening.py deleted file mode 100644 index 41a6b05..0000000 --- a/tests/test_security_hardening.py +++ /dev/null @@ -1,120 +0,0 @@ -from __future__ import annotations - -from pathlib import Path -import os - -import pytest -import yaml -from defusedxml.common import EntitiesForbidden - -from jinjaturtle import cli -from jinjaturtle.core import generate_ansible_yaml, parse_config, flatten_config -from jinjaturtle.multi import is_supported_file, iter_supported_files, process_directory -from jinjaturtle.output_safety import OutputPathError, write_text_safely - - -class UnsafeAwareLoader(yaml.SafeLoader): - pass - - -def _unsafe(loader: UnsafeAwareLoader, node: yaml.Node): - return loader.construct_scalar(node) - - -UnsafeAwareLoader.add_constructor("!unsafe", _unsafe) - - -def test_jinja_values_are_emitted_as_ansible_unsafe(tmp_path: Path): - src = tmp_path / "app.ini" - src.write_text("[main]\ncmd = {{ lookup('pipe','id') }}\n", encoding="utf-8") - - fmt, parsed = parse_config(src) - defaults_yaml = generate_ansible_yaml("role", flatten_config(fmt, parsed)) - - assert "role_main_cmd: !unsafe" in defaults_yaml - assert "{{ lookup(''pipe'',''id'') }}" in defaults_yaml - loaded = yaml.load(defaults_yaml, Loader=UnsafeAwareLoader) - assert loaded["role_main_cmd"] == "{{ lookup('pipe','id') }}" - - -def test_folder_mode_marks_nested_jinja_values_and_ids_unsafe(tmp_path: Path): - src = tmp_path / "src" - src.mkdir() - # Filename ids are source-derived values too. - (src / "{{ bad }}.yaml").write_text( - "message: \"{{ lookup('pipe','id') }}\"\n", encoding="utf-8" - ) - - defaults_yaml, _outputs = process_directory( - src, recursive=False, role_prefix="role" - ) - - assert "id: !unsafe" in defaults_yaml - assert "role_message: !unsafe" in defaults_yaml - loaded = yaml.load(defaults_yaml, Loader=UnsafeAwareLoader) - assert loaded["role_items"][0]["id"] == "{{ bad }}.yaml" - assert loaded["role_items"][0]["role_message"] == "{{ lookup('pipe','id') }}" - - -def test_xml_parser_rejects_entities_when_called_as_library(tmp_path: Path): - src = tmp_path / "bad.xml" - src.write_text( - "]>&xxe;", - encoding="utf-8", - ) - - with pytest.raises(EntitiesForbidden): - parse_config(src, "xml") - - -def test_folder_mode_does_not_follow_symlinked_files(tmp_path: Path): - real = tmp_path / "secret.ini" - real.write_text("[main]\nsecret=yes\n", encoding="utf-8") - root = tmp_path / "root" - root.mkdir() - link = root / "link.ini" - link.symlink_to(real) - - assert not is_supported_file(link) - assert iter_supported_files(root, recursive=False) == [] - assert iter_supported_files(root, recursive=True) == [] - - -@pytest.mark.skipif(not hasattr(os, "symlink"), reason="symlinks unavailable") -def test_cli_refuses_to_write_through_final_symlink(tmp_path: Path): - target = tmp_path / "target.txt" - target.write_text("keep\n", encoding="utf-8") - link = tmp_path / "out.yml" - link.symlink_to(target) - - with pytest.raises(OutputPathError): - write_text_safely(link, "replace\n") - - assert target.read_text(encoding="utf-8") == "keep\n" - assert link.is_symlink() - - -def test_cli_refuses_symlinked_output_parent(tmp_path: Path): - real_dir = tmp_path / "real" - real_dir.mkdir() - link_dir = tmp_path / "linkdir" - link_dir.symlink_to(real_dir, target_is_directory=True) - - with pytest.raises(OutputPathError): - write_text_safely(link_dir / "out.yml", "data\n") - - assert not (real_dir / "out.yml").exists() - - -def test_cli_reports_unsafe_output_path_without_overwriting_symlink(tmp_path: Path): - cfg = tmp_path / "app.ini" - cfg.write_text("[main]\nname = ok\n", encoding="utf-8") - target = tmp_path / "target.yml" - target.write_text("keep\n", encoding="utf-8") - link = tmp_path / "defaults.yml" - link.symlink_to(target) - - exit_code = cli._main([str(cfg), "--defaults-output", str(link)]) - - assert exit_code == 2 - assert target.read_text(encoding="utf-8") == "keep\n" diff --git a/tests/test_yaml_handler.py b/tests/test_yaml_handler.py index 548f5a6..543c759 100644 --- a/tests/test_yaml_handler.py +++ b/tests/test_yaml_handler.py @@ -177,17 +177,15 @@ def test_yaml_loop_preserves_blank_separator_after_list(tmp_path: Path): "blank", "require:\n - rubocop-performance\n - rubocop-rspec\n\nAllCops:\n NewCops: enable\n", "\n - rubocop-rspec\n\nAllCops:", - "<% end %>\nAllCops:", ), ( "no_blank", "require:\n - rubocop-performance\n - rubocop-rspec\nAllCops:\n NewCops: enable\n", "\n - rubocop-rspec\nAllCops:", - "<% end %>AllCops:", ), ] - for label, text, rendered_expected, erb_expected in cases: + for label, text, rendered_expected in cases: path = tmp_path / f"{label}.yml" path.write_text(text, encoding="utf-8")