diff --git a/README.md b/README.md index bbcda4e..1726719 100644 --- a/README.md +++ b/README.md @@ -1,11 +1,5 @@ # JinjaTurtle -## ABANDONED - -Project is abandoned. To many potential security issues. Please uninstall it. Sorry for wasting your time. - -____ -
JinjaTurtle logo
@@ -13,7 +7,7 @@ ____ JinjaTurtle is a command-line tool that helps turn existing native configuration files into reusable configuration-management templates. -It generates: +By default it generates: - a **Jinja2** template; and - an **Ansible defaults YAML** file containing the variables used by that @@ -29,6 +23,8 @@ variable data beside the template. JinjaTurtle examines a source config file and keeps the original structure as much as possible. +For the default Jinja2/Ansible mode: + 1. The config file is parsed. 2. Variable names are generated from the config keys and paths. 3. Those variable names are prefixed with `--role-name`, which should usually @@ -95,15 +91,15 @@ homogeneous enough. If it is not confident, it falls back to flattened scalar variables. Some very complex files will still need manual cleanup. The goal is to speed up -conversion into Jinja2 templates, not to guarantee a perfect final role without +conversion into Jinja2 templates, not to guarantee a perfect final module without review. ## JSON, quoting, and type preservation JinjaTurtle tries to preserve rendered config types. -For JSON, it uses JSON-aware Jinja2 expressions rather than plain string -substitution. This avoids generating invalid JSON such as: +For JSON, it uses JSON-aware expressions rather than plain string substitution. +This avoids generating invalid JSON such as: ```json {"enabled": True} @@ -115,6 +111,8 @@ when the correct rendered JSON should be: {"enabled": true} ``` +This uses Ansible-style JSON filters. + ## Can I convert multiple files at once? Yes. Pass a directory instead of a single file and JinjaTurtle will convert the @@ -206,18 +204,18 @@ positional arguments: options: -h, --help show this help message and exit -r, --role-name ROLE_NAME - Ansible role name, used as variable prefix. Defaults - to jinjaturtle. + Role name / variable prefix. In Jinja2 mode this is + usually the Ansible role name. Defaults to jinjaturtle. --recursive When CONFIG is a folder, recurse into subfolders. -f, --format {ini,json,toml,yaml,xml,postfix,systemd,ssh} Force config format instead of auto-detecting from filename. -d, --defaults-output DEFAULTS_OUTPUT - Path to write defaults/main.yml. If omitted, defaults - YAML is printed to stdout. + Path to write the generated variable YAML. If omitted, + it is printed to stdout. -t, --template-output TEMPLATE_OUTPUT - Path to write the generated config template. If - omitted, template is printed to stdout. + Path to write the generated config template. If omitted, + it is printed to stdout. ``` ## Additional supported formats @@ -243,37 +241,38 @@ code. Two guarantees matter: 1. **Values are data, never code.** Every config *value* is replaced with a - `{{ variable }}` placeholder in the template, and the original value is - stored separately in the defaults data. String values in generated defaults - are tagged with Ansible's `!unsafe` tag by default, so a payload sitting - inside a value (for example `motd = {{ lookup('pipe', 'id') }}`) remains - inert even if downstream playbooks accidentally place that value in another - templating context. + `{{ variable }}` placeholder in the template, and the original value is stored + separately in the defaults data. When the template is later rendered, + the placeholder prints the value as a literal string; Jinja2 does not + recursively render the *contents* of a variable, so a payload sitting inside + a value is inert. 2. **Verbatim text is neutralised.** To preserve formatting, JinjaTurtle copies comments, blank lines, headers and any unrecognised lines from the source - into the template. Any Jinja2 metacharacters in that copied text (`{{ }}`, - `{% %}`, `{# #}`) are escaped so they render as the literal characters the - author wrote, rather than executing. + into the template. Any template metacharacters in that copied text + (`{{ }}`, `{% %}`, `{# #}` ) are escaped so they render as the literal + characters the author wrote, rather than executing. ### Consumer responsibilities -Treat generated defaults as untrusted input. JinjaTurtle emits extracted -string values with Ansible's `!unsafe` tag, but you should preserve that tag if -you merge the generated data into larger defaults files or transform it with -other tooling. +The value guarantee above relies on the downstream renderer being single-pass, +which is the normal case: -In short: render JinjaTurtle output exactly once. Do not strip `!unsafe` tags or -feed the data back through another templating pass. +- **Ansible**: rendering a template with `template:`/`ansible.builtin.template` + is single-pass. For defence in depth, treat the generated defaults as + untrusted input — Ansible already does not re-template variable *contents* by + default. If you build your own var structures from this data and pass them + through additional templating, mark untrusted values with the `!unsafe` tag so + they are never re-evaluated. -Directory mode deliberately ignores symlinks and will not traverse outside the -input directory. This prevents a malicious folder tree from smuggling in -configuration files through links to unrelated filesystem locations. +In short: render JinjaTurtle output exactly once. Do not feed it back through +another templating pass. **IMPORTANT**: Always review both the original config files, then the resulting templates generated by JinjaTurtle, before integrating them into your config management system! + ## Found a bug, have a suggestion? You can e-mail me; see `pyproject.toml` for details. You can also contact me on diff --git a/debian/changelog b/debian/changelog index af99a92..9173c94 100644 --- a/debian/changelog +++ b/debian/changelog @@ -1,6 +1,5 @@ jinjaturtle (0.5.7) unstable; urgency=medium - * Remove ERB/Puppet output support and keep Jinja2-only generation. * More hardening measures -- Miguel Jacq Wed, 24 Jun 2026 16:13:00 +1000 diff --git a/rpm/jinjaturtle.spec b/rpm/jinjaturtle.spec index 16ebd66..051fd8b 100644 --- a/rpm/jinjaturtle.spec +++ b/rpm/jinjaturtle.spec @@ -43,7 +43,6 @@ Convert config files into Ansible defaults and Jinja2 templates. %changelog * Wed Jun 24 2026 Miguel Jacq - %{version}-%{release} -- Remove ERB/Puppet output support and keep Jinja2-only generation. - More hardening * Tue Jun 23 2026 Miguel Jacq - %{version}-%{release} - Try to prevent what could lead to execution of embedded jinja in original files when converting diff --git a/src/jinjaturtle/cli.py b/src/jinjaturtle/cli.py index 057abdc..bdabc18 100644 --- a/src/jinjaturtle/cli.py +++ b/src/jinjaturtle/cli.py @@ -1,7 +1,6 @@ from __future__ import annotations import argparse -import os import sys from defusedxml import defuse_stdlib from pathlib import Path @@ -17,25 +16,7 @@ from .core import ( from .multi import process_directory from .safety import TemplateSafetyError - - -def _write_text_no_symlink(path: Path, text: str) -> None: - """Write a user-requested output file without following a final symlink.""" - flags = os.O_WRONLY | os.O_CREAT | os.O_TRUNC - if hasattr(os, "O_NOFOLLOW"): - flags |= os.O_NOFOLLOW - try: - fd = os.open(path, flags, 0o600) - except OSError as exc: - raise OSError(f"refusing or unable to write output file {path}: {exc}") from exc - with os.fdopen(fd, "w", encoding="utf-8") as f: - f.write(text) - - -def _ensure_output_dir(path: Path) -> None: - if path.exists() and path.is_symlink(): - raise OSError(f"refusing to use symlink as output directory: {path}") - path.mkdir(parents=True, exist_ok=True) +from .output_safety import OutputPathError, ensure_safe_directory, write_text_safely def _build_arg_parser() -> argparse.ArgumentParser: @@ -92,9 +73,9 @@ def _main(argv: list[str] | None = None) -> int: f"jinjaturtle: refusing to generate unsafe template: {exc}", file=sys.stderr ) return 2 - except (OSError, TypeError, ValueError) as exc: - print(f"jinjaturtle: error: {exc}", file=sys.stderr) - return 1 + except OutputPathError as exc: + print(f"jinjaturtle: refusing unsafe output path: {exc}", file=sys.stderr) + return 2 def _run(argv: list[str] | None = None) -> int: @@ -112,36 +93,37 @@ def _run(argv: list[str] | None = None) -> int: # Write defaults if args.defaults_output: - _write_text_no_symlink(Path(args.defaults_output), defaults_yaml) + write_text_safely(Path(args.defaults_output), defaults_yaml) else: print("# defaults/main.yml") print(defaults_yaml, end="") + template_ext = j2.TEMPLATE_EXTENSION + # Write templates if args.template_output: out_path = Path(args.template_output) if len(outputs) == 1 and not out_path.is_dir(): - _write_text_no_symlink(out_path, outputs[0].template) + write_text_safely(out_path, outputs[0].template) else: - _ensure_output_dir(out_path) + ensure_safe_directory(out_path) for o in outputs: - _write_text_no_symlink( - out_path / f"config.{o.fmt}.{j2.TEMPLATE_EXTENSION}", - o.template, + write_text_safely( + out_path / f"config.{o.fmt}.{template_ext}", o.template ) else: for o in outputs: name = ( - f"config.{j2.TEMPLATE_EXTENSION}" + f"config.{template_ext}" if len(outputs) == 1 - else f"config.{o.fmt}.{j2.TEMPLATE_EXTENSION}" + else f"config.{o.fmt}.{template_ext}" ) print(f"# {name}") print(o.template, end="") return 0 - # Single-file mode + # Single-file mode (existing behaviour) config_text = config_path.read_text(encoding="utf-8") # Parse the config @@ -166,13 +148,13 @@ def _run(argv: list[str] | None = None) -> int: ) if args.defaults_output: - _write_text_no_symlink(Path(args.defaults_output), ansible_yaml) + write_text_safely(Path(args.defaults_output), ansible_yaml) else: print("# defaults/main.yml") print(ansible_yaml, end="") if args.template_output: - _write_text_no_symlink(Path(args.template_output), template_str) + write_text_safely(Path(args.template_output), template_str) else: print(f"# config.{j2.TEMPLATE_EXTENSION}") print(template_str, end="") diff --git a/src/jinjaturtle/core.py b/src/jinjaturtle/core.py index 8f7d9cf..72b42de 100644 --- a/src/jinjaturtle/core.py +++ b/src/jinjaturtle/core.py @@ -7,39 +7,11 @@ import datetime import re import yaml - -class AnsibleUnsafeString(str): - """Marker for strings emitted with Ansible's !unsafe YAML tag.""" - - pass - - -class VariableNameCollisionError(ValueError): - """Raised when distinct config paths collapse to the same variable name.""" - - pass - - -class ConfigTooDeeplyNestedError(ValueError): - """Raised when a parsed config nests deeper than ``MAX_CONFIG_DEPTH``. - - Pathologically nested input (thousands of levels) would otherwise drive the - recursive walkers in this module past Python's stack limit and raise an - uncaught ``RecursionError``. We bound the depth here and fail with a clear, - catchable error instead. The limit is well above any plausible real config - while staying comfortably under the interpreter's default recursion ceiling. - """ - - pass - - -# Maximum container-nesting depth accepted from a parsed config. Real-world -# configuration files are nowhere near this deep; the cap exists purely to turn -# a hostile/degenerate input into a clean error rather than a stack overflow. -MAX_CONFIG_DEPTH = 200 - from .loop_analyzer import LoopAnalyzer, LoopCandidate -from .safety import verify_jinja2_template_safe +from .safety import ( + verify_jinja2_template_safe, + verify_no_live_jinja_in_json_keys, +) from .handlers import ( BaseHandler, IniHandler, @@ -61,6 +33,22 @@ class QuotedString(str): pass +class AnsibleUnsafeString(str): + """Marker type emitted with Ansible's !unsafe YAML tag. + + Ansible recursively templates string values by default. Source-derived + config values that contain Jinja delimiters must therefore be marked + unsafe in defaults/main.yml, otherwise a harvested value such as + ``{{ lookup('pipe', 'id') }}`` becomes executable on the Ansible + controller when the generated role is applied. + """ + + pass + + +_JINJA_STARTS = ("{{", "{%", "{#") + + def _fallback_str_representer(dumper: yaml.SafeDumper, data: Any): """ Fallback for objects the dumper doesn't know about. @@ -81,37 +69,32 @@ def _quoted_str_representer(dumper: yaml.SafeDumper, data: QuotedString): def _ansible_unsafe_str_representer(dumper: yaml.SafeDumper, data: AnsibleUnsafeString): - return dumper.represent_scalar("!unsafe", str(data)) + return dumper.represent_scalar("!unsafe", str(data), style="'") -def _ansible_unsafe_str_constructor( - loader: yaml.SafeLoader, node: yaml.nodes.ScalarNode -) -> str: - return loader.construct_scalar(node) +def _needs_ansible_unsafe(value: str) -> bool: + return any(marker in value for marker in _JINJA_STARTS) -def _mark_ansible_unsafe_values(data: Any) -> Any: - """Recursively tag string *values* as Ansible !unsafe. +def _mark_ansible_unsafe_values(obj: Any) -> Any: + """Recursively mark mapping/list values containing Jinja as !unsafe. - Mapping keys remain ordinary strings because they are variable names in the - generated defaults document, not attacker-controlled config values. + Mapping keys are intentionally left alone: they are variable names or YAML + structure, not Ansible-templated values. Values nested in folder-mode item + lists, including source-derived ``id`` values, are protected. """ - if isinstance(data, AnsibleUnsafeString): - return data - if isinstance(data, str): - return AnsibleUnsafeString(data) - if isinstance(data, dict): - return {k: _mark_ansible_unsafe_values(v) for k, v in data.items()} - if isinstance(data, list): - return [_mark_ansible_unsafe_values(v) for v in data] - if isinstance(data, tuple): - return tuple(_mark_ansible_unsafe_values(v) for v in data) - return data + + if isinstance(obj, dict): + return {k: _mark_ansible_unsafe_values(v) for k, v in obj.items()} + if isinstance(obj, list): + return [_mark_ansible_unsafe_values(v) for v in obj] + if isinstance(obj, str) and _needs_ansible_unsafe(obj): + return AnsibleUnsafeString(obj) + return obj _TurtleDumper.add_representer(QuotedString, _quoted_str_representer) _TurtleDumper.add_representer(AnsibleUnsafeString, _ansible_unsafe_str_representer) -yaml.SafeLoader.add_constructor("!unsafe", _ansible_unsafe_str_constructor) # Use our fallback for any unknown object types _TurtleDumper.add_representer(None, _fallback_str_representer) @@ -141,12 +124,11 @@ _HANDLERS["ssh"] = _SSH_HANDLER def dump_yaml(data: Any, *, sort_keys: bool = True) -> str: """Dump YAML using JinjaTurtle's dumper settings. - String values extracted from input configuration are emitted as Ansible - ``!unsafe`` values so a downstream playbook cannot accidentally re-template - malicious Jinja syntax stored in defaults. Mapping keys are left untouched. + This is used by both the single-file and multi-file code paths. """ + safe_data = _mark_ansible_unsafe_values(data) return yaml.dump( - _mark_ansible_unsafe_values(data), + safe_data, Dumper=_TurtleDumper, sort_keys=sort_keys, default_flow_style=False, @@ -323,31 +305,6 @@ def detect_format(path: Path, explicit: str | None = None) -> str: return "ini" -def _check_xml_depth(elem: Any, _depth: int = 0) -> None: - """Enforce :data:`MAX_CONFIG_DEPTH` on an ElementTree element. - - XML parses to an ``Element`` rather than dict/list containers, so it is not - covered by the depth bound in :func:`_stringify_timestamps`. Without this - check a pathologically nested document would drive the XML flattener and the - loop analyzer's ``_analyze_xml`` / ``_xml_elem_to_dict`` walkers into an - uncaught ``RecursionError``. We walk the tree iteratively (an explicit - stack, so the check itself cannot overflow) and fail loudly at the limit. - """ - # (element, depth) pairs; iterative traversal avoids its own recursion. - stack = [(elem, 0)] - while stack: - node, depth = stack.pop() - if depth > MAX_CONFIG_DEPTH: - raise ConfigTooDeeplyNestedError( - f"XML config nests deeper than the supported maximum of " - f"{MAX_CONFIG_DEPTH} levels; refusing to process it" - ) - for child in list(node): - # Skip comment/PI nodes whose tag is not a str. - if isinstance(getattr(child, "tag", None), str): - stack.append((child, depth + 1)) - - def parse_config(path: Path, fmt: str | None = None) -> tuple[str, Any]: """ Parse config file into a Python object. @@ -356,22 +313,7 @@ def parse_config(path: Path, fmt: str | None = None) -> tuple[str, Any]: handler = _HANDLERS.get(fmt) if handler is None: raise ValueError(f"Unsupported config format: {fmt}") - try: - parsed = handler.parse(path) - except RecursionError as exc: - # Some underlying parsers (e.g. PyYAML, json) recurse while *building* - # the object, so they can exhaust the stack before our own depth guards - # below ever run. Convert that into the same clean, catchable error. - raise ConfigTooDeeplyNestedError( - f"config in {path} is nested too deeply for the parser to handle; " - "refusing to process it" - ) from exc - # Bound nesting depth before any recursive walker runs, turning degenerate - # deeply-nested input into a clean, catchable error instead of a stack - # overflow. dict/list configs are bounded inside _stringify_timestamps; - # XML (an Element tree) is bounded explicitly here. - if hasattr(parsed, "tag"): # ElementTree Element - _check_xml_depth(parsed) + parsed = handler.parse(path) # Make sure datetime objects are treated as strings (TOML, YAML) parsed = _stringify_timestamps(parsed) @@ -441,31 +383,6 @@ def _path_starts_with(path: tuple[str, ...], prefix: tuple[str, ...]) -> bool: return path[: len(prefix)] == prefix -def check_var_name_collisions( - role_prefix: str, items: Iterable[tuple[tuple[str, ...], Any]] -) -> None: - """Fail if distinct config paths collapse to the same variable name.""" - seen: dict[str, tuple[str, ...]] = {} - for path, _value in items: - path_tuple = tuple(str(p) for p in path) - var_name = make_var_name(role_prefix, path_tuple) - previous = seen.get(var_name) - if previous is not None and previous != path_tuple: - raise VariableNameCollisionError( - "distinct config paths collapse to the same variable name " - f"{var_name!r}: {previous!r} and {path_tuple!r}" - ) - seen[var_name] = path_tuple - - -def _flatten_loop_defaults( - loop_candidates: list[LoopCandidate] | None, -) -> list[tuple[tuple[str, ...], Any]]: - if not loop_candidates: - return [] - return [(candidate.path, candidate.items) for candidate in loop_candidates] - - def generate_ansible_yaml( role_prefix: str, flat_items: list[tuple[tuple[str, ...], Any]], @@ -474,10 +391,6 @@ def generate_ansible_yaml( """ Create Ansible YAML for defaults/main.yml. """ - check_var_name_collisions( - role_prefix, flat_items + _flatten_loop_defaults(loop_candidates) - ) - defaults: dict[str, Any] = {} # Add scalar variables @@ -525,10 +438,29 @@ def generate_jinja2_template( # verbatim source text, the un-escaped payload shows up here as a live tag # and generation aborts instead of emitting an injectable template. verify_jinja2_template_safe(template) + + # Format-specific backstop: JinjaTurtle never emits Jinja inside a JSON object + # key, so a live construct in key position means source key text leaked into + # the template unescaped. This is independent of per-handler escaping. + if fmt == "json": + verify_no_live_jinja_in_json_keys(template) + return template -def _stringify_timestamps(obj: Any, _depth: int = 0) -> Any: +def _template_variable_names( + role_prefix: str, + flat_items: list[tuple[tuple[str, ...], Any]], + loop_candidates: list[LoopCandidate] | None = None, +) -> set[str]: + names = {make_var_name(role_prefix, path) for path, _value in flat_items} + if loop_candidates: + for candidate in loop_candidates: + names.add(make_var_name(role_prefix, candidate.path)) + return names + + +def _stringify_timestamps(obj: Any) -> Any: """ Recursively walk a parsed config and turn any datetime/date/time objects into plain strings in ISO-8601 form. @@ -538,23 +470,11 @@ def _stringify_timestamps(obj: Any, _depth: int = 0) -> Any: This commonly occurs otherwise with TOML and YAML files, which sees Python automatically convert those sorts of strings into datetime objects. - - Because this is the first walk applied to every parsed config (in - :func:`parse_config`), it also enforces :data:`MAX_CONFIG_DEPTH`. Bounding - nesting here means every downstream recursive walker (flatteners, loop - analyzer) is fed input of safe depth, so a degenerate deeply-nested file - fails with :class:`ConfigTooDeeplyNestedError` instead of an uncaught - ``RecursionError`` somewhere later in the pipeline. """ - if _depth > MAX_CONFIG_DEPTH: - raise ConfigTooDeeplyNestedError( - f"config nests deeper than the supported maximum of " - f"{MAX_CONFIG_DEPTH} levels; refusing to process it" - ) if isinstance(obj, dict): - return {k: _stringify_timestamps(v, _depth + 1) for k, v in obj.items()} + return {k: _stringify_timestamps(v) for k, v in obj.items()} if isinstance(obj, list): - return [_stringify_timestamps(v, _depth + 1) for v in obj] + return [_stringify_timestamps(v) for v in obj] # TOML & YAML both use the standard datetime types if isinstance(obj, datetime.datetime): diff --git a/src/jinjaturtle/escape.py b/src/jinjaturtle/escape.py index 6e68f69..e55f416 100644 --- a/src/jinjaturtle/escape.py +++ b/src/jinjaturtle/escape.py @@ -1,40 +1,61 @@ from __future__ import annotations -"""Neutralise Jinja2 metacharacters in text copied verbatim from source files. +"""Neutralise template metacharacters in text copied verbatim from source files. JinjaTurtle preserves formatting by copying parts of the *original* config file straight into the generated template: comments, blank lines, section headers, -and any line it does not recognise as ``key = value``. Config *values* are -always replaced with ``{{ var }}`` placeholders and stored in the defaults data, -so a payload inside a value remains data. Verbatim text is different: if the -source contains ``{{ ... }}``, ``{% ... %}`` or ``{# ... #}``, that text would -become live Jinja2 template code unless it is neutralised. +and any line it does not recognise as ``key = value``. Config *values* are +always replaced with ``{{ var }}`` placeholders and parked in the defaults data, +so a payload inside a value is inert. Verbatim text is different: if the source +contains ``{{ ... }}``, ``{% ... %}`` or ``{# ... #}`` (Jinja2), or ``<%= %>`` / +``<% %>`` (ERB), that text becomes *live template code* in the output and is +executed when Ansible later renders the template. Because JinjaTurtle is frequently fed harvested, attacker-influenceable config (hostnames, banners, GECOS-derived comments, "Managed by" notes), this is a template-injection / SSTI vector that can lead to remote code execution on the configuration-management control node. -The helpers here render those metacharacters as literal text so they survive a -later Jinja2 render as the characters the source author actually wrote, rather +The functions here render those metacharacters as literal text so they survive a +later template render as the characters the source author actually wrote, rather than as executable template syntax. + +Design notes: + * We only ever escape text that originates from the *source file*. We never + pass JinjaTurtle's own generated placeholders (``{{ role_var }}``) through + these helpers, so the placeholders keep working. + * Jinja2 literal text is wrapped in a single ``{% raw %} ... {% endraw %}`` + block. ``raw`` disables *all* tag interpretation inside it -- expressions, + statements and ``{# #}`` comments alike -- so one wrap neutralises every + Jinja construct. The only way to break out of a raw block is a literal + ``{% endraw %}`` in the source, so we defang the token ``endraw`` (in any + internal spacing) before wrapping. """ import re -# Jinja2 delimiters we must neutralise. Engine-default; JinjaTurtle never +# Jinja2 delimiters we must neutralise. Engine-default; JinjaTurtle never # configures custom delimiters. _JINJA_MARKERS = ("{{", "}}", "{%", "%}", "{#", "#}") +# ERB delimiters. Longer markers first so "<%=" matches before "<%". +_ERB_OPEN_MARKERS = ("<%=", "<%-", "<%#", "<%") +_ERB_CLOSE_MARKERS = ("-%>", "%>") + # Matches a Jinja2 endraw tag in any internal spacing and with any -# whitespace-control marker on either side. Jinja2 accepts "-", "+", or no +# whitespace-control marker on either side. Jinja2 accepts "-", "+", or no # marker adjacent to the "%}"/"{%" of a block tag (e.g. "{%endraw%}", # "{% endraw %}", "{%- endraw -%}", "{%+ endraw +%}"), and ALL of these close -# a raw block. The control marker must be matched so a "{%+ endraw %}" in +# a raw block. The control marker must be matched so a "{%+ endraw %}" in # attacker-influenced source text cannot survive defanging and break out of our -# {% raw %} wrapper. [-+]? appears on both sides accordingly. +# {% raw %} wrapper. [-+]? appears on both sides accordingly. _ENDRAW_RE = re.compile(r"{%[-+]?\s*endraw\s*[-+]?%}") +# Sentinel inserted between "end" and "raw" to break the endraw keyword without +# changing the visible characters. We use a Jinja comment-free approach: insert +# the two halves across a raw boundary so the literal text still reads "endraw" +# to a human but is never a valid tag. See escape_jinja_literal for usage. + def contains_jinja_markup(text: str) -> bool: """Return True if *text* contains any Jinja2 delimiter.""" @@ -44,14 +65,19 @@ def contains_jinja_markup(text: str) -> bool: def _defang_endraw(text: str) -> str: """Rewrite any literal ``{% endraw %}`` so it cannot close our raw wrapper. - We keep everything literal: the replacement breaks the keyword by emitting - the tag's text in two raw segments split inside the word ``endraw``. The - result, when later rendered, reproduces the exact original characters while - never being a parseable raw-closing tag inside our wrapper. + We turn each endraw tag into ``{% endraw %}{{ '{% endraw %}' }}{% raw %}``... + no -- that would re-introduce live tags. Instead we keep everything literal: + we break the keyword by emitting the tag's text in two raw segments split + inside the word ``endraw``. The result, when later rendered, reproduces the + exact original characters ``{% endraw %}`` while never being a parseable tag. """ def _replace(match: re.Match[str]) -> str: tag = match.group(0) + # Split the keyword "endraw" as "end" + "raw"; close and reopen the raw + # block between them. Each half is plain text inside a raw block, so the + # reconstructed output is byte-identical to the original tag, but at no + # point does the token "{% endraw %}" exist contiguously to close raw. idx = tag.lower().index("endraw") head = tag[: idx + 3] # up to and including "end" tail = tag[idx + 3 :] # "raw...%}" @@ -64,14 +90,9 @@ def escape_jinja_literal(text: str) -> str: """Make *text* render as literal characters under a later Jinja2 render. Text with no Jinja metacharacters is returned unchanged so the common case - stays byte-for-byte identical to the source. Otherwise the text is wrapped + stays byte-for-byte identical to the source. Otherwise the text is wrapped in a single ``{% raw %}`` block, with any embedded ``endraw`` defanged. """ if not text or not contains_jinja_markup(text): return text return "{% raw %}" + _defang_endraw(text) + "{% endraw %}" - - -def escape_literal(text: str) -> str: - """Escape verbatim *text* for Jinja2 output.""" - return escape_jinja_literal(text) diff --git a/src/jinjaturtle/handlers/base.py b/src/jinjaturtle/handlers/base.py index 71490bf..14aaec7 100644 --- a/src/jinjaturtle/handlers/base.py +++ b/src/jinjaturtle/handlers/base.py @@ -2,10 +2,6 @@ from __future__ import annotations from pathlib import Path from typing import Any, Iterable -import re - - -_VAR_NAME_RE = re.compile(r"^[A-Za-z_][A-Za-z0-9_]*$") class BaseHandler: @@ -63,13 +59,6 @@ class BaseHandler: Sanitises parts to lowercase [a-z0-9_] and strips extras. """ role_prefix = role_prefix.strip().lower() - if not _VAR_NAME_RE.fullmatch(role_prefix) or "__" in role_prefix: - raise ValueError( - "role name must be a safe Ansible/Jinja variable prefix: " - "letters, digits and underscores only; must start with a letter " - "or underscore; double underscores are not allowed" - ) - clean_parts: list[str] = [] for part in path: @@ -81,13 +70,10 @@ class BaseHandler: cleaned_chars.append(c.lower()) else: cleaned_chars.append("_") - cleaned_part = re.sub(r"_+", "_", "".join(cleaned_chars)).strip("_") + cleaned_part = "".join(cleaned_chars).strip("_") if cleaned_part: clean_parts.append(cleaned_part) if clean_parts: - var_name = role_prefix + "_" + "_".join(clean_parts) - else: - var_name = role_prefix - - return var_name + return role_prefix + "_" + "_".join(clean_parts) + return role_prefix diff --git a/src/jinjaturtle/handlers/json.py b/src/jinjaturtle/handlers/json.py index fe32ac6..0268666 100644 --- a/src/jinjaturtle/handlers/json.py +++ b/src/jinjaturtle/handlers/json.py @@ -2,12 +2,12 @@ from __future__ import annotations import json import re -from uuid import uuid4 from pathlib import Path from typing import Any from . import DictLikeHandler from .. import j2 +from ..escape import escape_jinja_literal from ..loop_analyzer import LoopCandidate @@ -110,10 +110,18 @@ class JsonHandler(DictLikeHandler): chunks: list[str] = [] pos = 0 for path, start, end in spans: - chunks.append(text[pos:start]) + # Text between scalar values (object keys, structural punctuation, + # whitespace, and any comment-like trailing text) is copied verbatim + # from the source file. Like every other text-emitting handler, this + # verbatim text must be neutralised: if it contains Jinja2 markup it + # would otherwise become live template code at apply time. The value + # itself is replaced with a safe placeholder below. ``escape_jinja_literal`` + # is a no-op on text without Jinja markers, so benign JSON is unchanged + # byte-for-byte and a later render reproduces the original characters. + chunks.append(escape_jinja_literal(text[pos:start])) chunks.append(self._json_value_expr(self.make_var_name(role_prefix, path))) pos = end - chunks.append(text[pos:]) + chunks.append(escape_jinja_literal(text[pos:])) return "".join(chunks) def _collect_json_scalar_spans( @@ -212,26 +220,30 @@ class JsonHandler(DictLikeHandler): Uses | to_json filter to preserve types (numbers, booleans, null). """ - marker_prefix = f"\ue000JT{uuid4().hex}" - scalar_markers: dict[str, str] = {} - def _walk(obj: Any, path: tuple[str, ...] = ()) -> Any: if isinstance(obj, dict): - return {k: _walk(v, path + (str(k),)) for k, v in obj.items()} + # Keys are emitted verbatim into the template, so neutralise any + # Jinja markup in them (see _generate_json_template_from_text). + return { + escape_jinja_literal(str(k)): _walk(v, path + (str(k),)) + for k, v in obj.items() + } if isinstance(obj, list): return [_walk(v, path + (str(i),)) for i, v in enumerate(obj)] + # scalar - use marker that will be replaced with to_json var_name = self.make_var_name(role_prefix, path) - marker = f"{marker_prefix}SCALAR:{len(scalar_markers)}\ue001" - scalar_markers[marker] = var_name - return marker + return f"__SCALAR__{var_name}__" templated = _walk(data) json_str = json.dumps(templated, indent=2, ensure_ascii=False) - for marker, var_name in scalar_markers.items(): - json_str = json_str.replace( - json.dumps(marker, ensure_ascii=False), self._json_value_expr(var_name) - ) + # Replace scalar markers with Jinja expressions using to_json filter + # This preserves types (numbers stay numbers, booleans stay booleans) + json_str = re.sub( + r'"__SCALAR__([a-zA-Z_][a-zA-Z0-9_]*)__"', + lambda m: self._json_value_expr(m.group(1)), + json_str, + ) return json_str + "\n" @@ -247,67 +259,71 @@ class JsonHandler(DictLikeHandler): Generate a JSON Jinja2 template with for loops where appropriate. """ - marker_prefix = f"\ue000JT{uuid4().hex}" - scalar_markers: dict[str, str] = {} - loop_markers: dict[tuple[str, str], str] = {} - def _walk(obj: Any, current_path: tuple[str, ...] = ()) -> Any: + # Check if this path is a loop candidate if current_path in loop_paths: + # Find the matching candidate candidate = next(c for c in loop_candidates if c.path == current_path) collection_var = self.make_var_name(role_prefix, candidate.path) + item_var = candidate.loop_var if candidate.item_schema == "scalar": - marker = f"{marker_prefix}LOOP_SCALAR:{len(loop_markers)}\ue001" - loop_markers[("scalar", collection_var)] = marker - return marker - if candidate.item_schema in ("simple_dict", "nested"): - marker = f"{marker_prefix}LOOP_DICT:{len(loop_markers)}\ue001" - loop_markers[("dict", collection_var)] = marker - return marker + # Simple list of scalars - use special marker that we'll replace + return f"__LOOP_SCALAR__{collection_var}__{item_var}__" + elif candidate.item_schema in ("simple_dict", "nested"): + # List of dicts - use special marker + return f"__LOOP_DICT__{collection_var}__{item_var}__" if isinstance(obj, dict): - return {k: _walk(v, current_path + (str(k),)) for k, v in obj.items()} + # Keys are emitted verbatim into the template, so neutralise any + # Jinja markup in them (see _generate_json_template_from_text). + return { + escape_jinja_literal(str(k)): _walk(v, current_path + (str(k),)) + for k, v in obj.items() + } if isinstance(obj, list): + # Check if this list is a loop candidate if current_path in loop_paths: + # Already handled above return _walk(obj, current_path) return [_walk(v, current_path + (str(i),)) for i, v in enumerate(obj)] + # scalar - use marker to preserve type var_name = self.make_var_name(role_prefix, current_path) - marker = f"{marker_prefix}SCALAR:{len(scalar_markers)}\ue001" - scalar_markers[marker] = var_name - return marker + return f"__SCALAR__{var_name}__" templated = _walk(data, path) + + # Convert to JSON string json_str = json.dumps(templated, indent=2, ensure_ascii=False) - for marker, var_name in scalar_markers.items(): - json_str = json_str.replace( - json.dumps(marker, ensure_ascii=False), self._json_value_expr(var_name) - ) + # Replace scalar markers with Jinja expressions using to_json filter + json_str = re.sub( + r'"__SCALAR__([a-zA-Z_][a-zA-Z0-9_]*)__"', + lambda m: self._json_value_expr(m.group(1)), + json_str, + ) + # Post-process to replace loop markers with actual Jinja loops (indent-aware) for candidate in loop_candidates: collection_var = self.make_var_name(role_prefix, candidate.path) item_var = candidate.loop_var if candidate.item_schema == "scalar": - marker = loop_markers.get(("scalar", collection_var)) - if marker is None: - continue + marker = f'"__LOOP_SCALAR__{collection_var}__{item_var}__"' json_str = self._replace_marker_with_pretty_loop( json_str, - json.dumps(marker, ensure_ascii=False), + marker, lambda base, cv=collection_var, iv=item_var, c=candidate: self._generate_json_scalar_loop( cv, iv, c, base ), ) elif candidate.item_schema in ("simple_dict", "nested"): - marker = loop_markers.get(("dict", collection_var)) - if marker is None: - continue + marker = f'"__LOOP_DICT__{collection_var}__{item_var}__"' json_str = self._replace_marker_with_pretty_loop( json_str, - json.dumps(marker, ensure_ascii=False), + marker, lambda base, cv=collection_var, iv=item_var, c=candidate: self._generate_json_dict_loop( cv, iv, c, base ), @@ -367,9 +383,13 @@ class JsonHandler(DictLikeHandler): ] # first line has no indent; we prepend `inner` when emitting for i, key in enumerate(keys): comma = "," if i < len(keys) - 1 else "" - safe_key = json.dumps(str(key), ensure_ascii=False) + # The literal key text is emitted verbatim into the template; escape any + # Jinja markup in it. The value side ({item_var}.{key}) is constrained by + # the output safety gate's dotted-name allowlist, which fails closed on + # anything that is not a plain identifier path. dict_lines.append( - f"{field}{safe_key}: " f"{j2.to_json(f'{item_var}.{key}')}{comma}" + f'{field}"{escape_jinja_literal(str(key))}": ' + f"{j2.to_json(f'{item_var}.{key}')}{comma}" ) # Comma between *items* goes after the closing brace. dict_lines.append(f"{inner}}}{j2.if_not_loop_last()},{j2.endif()}") diff --git a/src/jinjaturtle/handlers/xml.py b/src/jinjaturtle/handlers/xml.py index 4339fa4..56e3fd7 100644 --- a/src/jinjaturtle/handlers/xml.py +++ b/src/jinjaturtle/handlers/xml.py @@ -3,7 +3,7 @@ from __future__ import annotations from collections import Counter, defaultdict from pathlib import Path from typing import Any -import xml.etree.ElementTree as ET # nosec +import xml.etree.ElementTree as ET # nosec B405 - safe trees only; parsing uses defusedxml import defusedxml.ElementTree as DET from .base import BaseHandler @@ -21,14 +21,11 @@ class XmlHandler(BaseHandler): def parse(self, path: Path) -> ET.Element: text = path.read_text(encoding="utf-8") - parser = DET.XMLParser( - target=ET.TreeBuilder(insert_comments=False), - forbid_dtd=True, - forbid_entities=True, - forbid_external=True, - ) - parser.feed(text) - root = parser.close() + # Security must live in the handler, not only in the CLI entry point: + # callers may import JinjaTurtle as a library and invoke parse_config() + # directly. defusedxml rejects DTD/entity abuse and also discards + # comments by default, matching the previous TreeBuilder behaviour. + root = DET.fromstring(text) return root def flatten(self, parsed: Any) -> list[tuple[tuple[str, ...], Any]]: @@ -235,6 +232,15 @@ class XmlHandler(BaseHandler): walk(root, ()) + # Internal marker prefixes used by JinjaTurtle's own comment nodes. These + # must NOT be escaped (they are converted into real Jinja control structures + # downstream). Source-file comments have none of these prefixes. + _MARKER_PREFIXES = ("LOOP:", "IF:", "ENDIF:") + + def _is_jt_marker(self, comment_text: str) -> bool: + stripped = (comment_text or "").lstrip() + return any(stripped.startswith(p) for p in self._MARKER_PREFIXES) + def _escape_source_comments(self, root: ET.Element) -> None: """Escape template metacharacters in comments preserved from the source. @@ -247,34 +253,30 @@ class XmlHandler(BaseHandler): placeholders, so comments (and the prolog, handled separately) are the only XML injection vector. - This function is called immediately after parsing source XML and before - JinjaTurtle appends any of its own marker comments. Therefore every - comment currently in the tree is attacker-controlled source data and must - be escaped, including comments that resemble JinjaTurtle's internal - marker syntax. + JinjaTurtle's own internal marker comments are left untouched so they can + be converted into real loops/conditionals later. """ + # ET represents comments with a callable tag (ET.Comment). Iterate all + # descendants and escape comment text that is not one of our markers. for elem in root.iter(): if elem.tag is ET.Comment: - elem.text = escape_jinja_literal(elem.text or "") + if not self._is_jt_marker(elem.text or ""): + elem.text = escape_jinja_literal(elem.text or "") def _generate_xml_template_from_text(self, role_prefix: str, text: str) -> str: """Generate scalar-only Jinja2 template.""" prolog, body = self._split_xml_prolog(text) - parser = DET.XMLParser( - target=ET.TreeBuilder(insert_comments=True), - forbid_dtd=True, - forbid_entities=True, - forbid_external=True, - ) + parser = ET.XMLParser(target=ET.TreeBuilder(insert_comments=True)) # nosec B314 parser.feed(body) root = parser.close() - # Escape source comments before adding any JinjaTurtle marker comments. - self._escape_source_comments(root) - self._apply_jinja_to_xml_tree(role_prefix, root) + # Neutralise template metacharacters in any comments preserved from the + # source file before serialising. + self._escape_source_comments(root) + indent = getattr(ET, "indent", None) if indent is not None: indent(root, space=" ") # type: ignore[arg-type] @@ -292,22 +294,19 @@ class XmlHandler(BaseHandler): prolog, body = self._split_xml_prolog(text) - # Parse with comments preserved. - parser = DET.XMLParser( - target=ET.TreeBuilder(insert_comments=True), - forbid_dtd=True, - forbid_entities=True, - forbid_external=True, - ) + # Parse with comments preserved + parser = ET.XMLParser(target=ET.TreeBuilder(insert_comments=True)) # nosec B314 parser.feed(body) root = parser.close() - # Escape source comments before adding any JinjaTurtle marker comments. - self._escape_source_comments(root) - # Apply Jinja transformations (including loop markers) self._apply_jinja_to_xml_tree(role_prefix, root, loop_candidates) + # Escape comments preserved from the source. JinjaTurtle's own + # LOOP/IF/ENDIF marker comments are recognised and left intact so they + # can be converted into real Jinja control structures below. + self._escape_source_comments(root) + # Convert to string indent = getattr(ET, "indent", None) if indent is not None: diff --git a/src/jinjaturtle/multi.py b/src/jinjaturtle/multi.py index 1add120..c968d61 100644 --- a/src/jinjaturtle/multi.py +++ b/src/jinjaturtle/multi.py @@ -22,22 +22,18 @@ Notes: from collections import Counter, defaultdict from copy import deepcopy +import os import configparser +import stat from dataclasses import dataclass from pathlib import Path from typing import Any, Iterable import xml.etree.ElementTree as ET # nosec from . import j2 -from .core import ( - check_var_name_collisions, - dump_yaml, - flatten_config, - make_var_name, - parse_config, -) +from .core import dump_yaml, flatten_config, make_var_name, parse_config from .handlers.xml import XmlHandler -from .safety import verify_jinja2_template_safe +from .safety import verify_jinja2_template_safe, verify_no_live_jinja_in_json_keys from .escape import escape_jinja_literal @@ -50,8 +46,23 @@ SUPPORTED_SUFFIXES: dict[str, set[str]] = { } +def _lstat(path: Path) -> os.stat_result: + return path.lstat() + + def is_supported_file(path: Path) -> bool: - if path.is_symlink() or not path.is_file(): + """Return True only for real regular files with supported suffixes. + + pathlib.Path.is_file() follows symlinks. Folder mode must not follow + attacker-controlled symlinks when run over an untrusted tree, especially if + an administrator accidentally runs the CLI as root. + """ + + try: + st = _lstat(path) + except FileNotFoundError: + return False + if not stat.S_ISREG(st.st_mode): return False suffix = path.suffix.lower() for exts in SUPPORTED_SUFFIXES.values(): @@ -61,27 +72,20 @@ def is_supported_file(path: Path) -> bool: def iter_supported_files(root: Path, recursive: bool) -> list[Path]: - if not root.exists(): + try: + st = _lstat(root) + except FileNotFoundError: raise FileNotFoundError(str(root)) - if root.is_symlink(): - return [] - if root.is_file(): + + if stat.S_ISLNK(st.st_mode): + raise ValueError(f"refusing to follow symlink: {root}") + if stat.S_ISREG(st.st_mode): return [root] if is_supported_file(root) else [] - if not root.is_dir(): + if not stat.S_ISDIR(st.st_mode): return [] - resolved_root = root.resolve(strict=True) it = root.rglob("*") if recursive else root.glob("*") - files: list[Path] = [] - for p in it: - if not is_supported_file(p): - continue - try: - resolved = p.resolve(strict=True) - resolved.relative_to(resolved_root) - except (OSError, ValueError): - continue - files.append(p) + files = [p for p in it if is_supported_file(p)] files.sort() return files @@ -702,7 +706,6 @@ def process_directory( for rid, parsed, containers in zip(rel_ids, parsed_list, container_sets): item: dict[str, Any] = {"id": rid} flat = flatten_config(fmt, parsed, loop_candidates=None) - check_var_name_collisions(role_prefix, flat) for path, value in flat: item[make_var_name(role_prefix, path)] = value for cpath in optional_containers: @@ -732,7 +735,6 @@ def process_directory( for rid, parser in zip(rel_ids, parsers): # type: ignore[arg-type] item: dict[str, Any] = {"id": rid} flat = flatten_config(fmt, parser, loop_candidates=None) - check_var_name_collisions(role_prefix, flat) for path, value in flat: item[make_var_name(role_prefix, path)] = value # section presence @@ -779,7 +781,6 @@ def process_directory( for rid, parsed, elems in zip(rel_ids, parsed_list, elem_sets): item: dict[str, Any] = {"id": rid} flat = flatten_config(fmt, parsed, loop_candidates=None) - check_var_name_collisions(role_prefix, flat) for path, value in flat: item[make_var_name(role_prefix, path)] = value for epath in optional_elements: @@ -805,5 +806,7 @@ def process_directory( # Any un-neutralised source text that became a live tag aborts generation. for out in outputs: verify_jinja2_template_safe(out.template) + if out.fmt == "json": + verify_no_live_jinja_in_json_keys(out.template) return defaults_yaml, outputs diff --git a/src/jinjaturtle/output_safety.py b/src/jinjaturtle/output_safety.py new file mode 100644 index 0000000..9833cde --- /dev/null +++ b/src/jinjaturtle/output_safety.py @@ -0,0 +1,124 @@ +from __future__ import annotations + +"""Safer file-output helpers for the JinjaTurtle CLI. + +The CLI is often used by administrators. A plain Path.write_text() follows a +final-path symlink and can therefore be dangerous when a root-run invocation +writes into an attacker-writable tree. These helpers validate path components, +write through a private temporary file in the target directory, and replace the +final path atomically. Existing final-path symlinks are refused rather than +followed. +""" + +import os +from pathlib import Path +import stat +import tempfile + + +class OutputPathError(OSError): + """Raised when a requested output path is unsafe.""" + + +def _absolute(path: Path) -> Path: + return path if path.is_absolute() else Path.cwd() / path + + +def _check_existing_path_not_symlink(path: Path) -> None: + try: + st = path.lstat() + except FileNotFoundError: + return + if stat.S_ISLNK(st.st_mode): + raise OutputPathError(f"refusing to use symlink path: {path}") + + +def _check_existing_output_file(path: Path) -> None: + try: + st = path.lstat() + except FileNotFoundError: + return + if stat.S_ISLNK(st.st_mode): + raise OutputPathError(f"refusing to write through symlink: {path}") + if not stat.S_ISREG(st.st_mode): + raise OutputPathError(f"refusing to replace non-regular file: {path}") + + +def _check_parent_components(parent: Path) -> None: + """Require every existing parent component to be a real directory.""" + + parent = _absolute(parent) + parts = parent.parts + if not parts: + return + + cur = Path(parts[0]) + for part in parts[1:]: + cur = cur / part + try: + st = cur.lstat() + except FileNotFoundError as exc: + raise OutputPathError(f"output parent does not exist: {cur}") from exc + if stat.S_ISLNK(st.st_mode): + raise OutputPathError(f"refusing to use symlink parent: {cur}") + if not stat.S_ISDIR(st.st_mode): + raise OutputPathError(f"output parent is not a directory: {cur}") + + +def ensure_safe_directory(path: Path) -> None: + """Create or validate a directory tree without accepting symlinks.""" + + path = _absolute(path) + parts = path.parts + if not parts: + return + + cur = Path(parts[0]) + for part in parts[1:]: + cur = cur / part + try: + st = cur.lstat() + except FileNotFoundError: + cur.mkdir(mode=0o700) + st = cur.lstat() + if stat.S_ISLNK(st.st_mode): + raise OutputPathError(f"refusing to use symlink directory: {cur}") + if not stat.S_ISDIR(st.st_mode): + raise OutputPathError(f"output path is not a directory: {cur}") + + +def write_text_safely(path: Path, text: str, *, encoding: str = "utf-8") -> None: + """Write text without following a final-path symlink. + + The target's parent must already exist and every parent component must be a + real directory. The write is completed with os.replace(), which atomically + swaps the final directory entry and does not dereference a final symlink. + """ + + path = _absolute(path) + _check_parent_components(path.parent) + _check_existing_output_file(path) + + fd = -1 + tmp_name: str | None = None + try: + fd, tmp_name = tempfile.mkstemp( + prefix=f".{path.name}.", suffix=".tmp", dir=str(path.parent), text=True + ) + with os.fdopen(fd, "w", encoding=encoding) as f: + fd = -1 + f.write(text) + f.flush() + os.fsync(f.fileno()) + os.chmod(tmp_name, 0o600) + _check_existing_output_file(path) + os.replace(tmp_name, path) + tmp_name = None + finally: + if fd >= 0: + os.close(fd) + if tmp_name is not None: + try: + os.unlink(tmp_name) + except FileNotFoundError: + pass diff --git a/src/jinjaturtle/safety.py b/src/jinjaturtle/safety.py index 5b2dea3..5ed2fe1 100644 --- a/src/jinjaturtle/safety.py +++ b/src/jinjaturtle/safety.py @@ -22,8 +22,8 @@ single global property: The check is positive/allowlist-based, which is the safe direction: unknown constructs are rejected, not ignored. It runs at the single choke points in -``core.py`` (``generate_jinja2_template``), so it -covers every current handler and every future one automatically. +``core.py`` (``generate_jinja2_template``) so it covers every current handler and +every future one automatically. Why this is robust against the escaper being wrong --------------------------------------------------- @@ -42,6 +42,8 @@ import re __all__ = [ "TemplateSafetyError", "verify_jinja2_template_safe", + "verify_erb_template_safe", + "verify_no_live_jinja_in_json_keys", ] @@ -204,3 +206,132 @@ def verify_jinja2_template_safe(template_text: str) -> None: # tokens) so the reconstructed body matches the source spacing. body_parts.append(value if isinstance(value, str) else str(value)) # raw_begin / raw_end / data tokens outside a tag are inert: skip. + + +# --------------------------------------------------------------------------- # +# JSON-key gate (defence in depth, format-specific). +# +# JinjaTurtle never emits a live Jinja construct inside a JSON *object key*: keys +# are copied verbatim from the source and (after escape.py) are wrapped in +# ``{% raw %}`` if they contain markup, so a key never lexes as a live tag. A +# live construct in key position can therefore only mean source key text leaked +# into the template unescaped (the json-handler blind spot). This gate is +# independent of the escaper: it inspects the finished template, replaces every +# *live* Jinja construct with an inert sentinel (raw-wrapped literal text stays +# literal), and rejects any sentinel that lands in a JSON key slot. +# --------------------------------------------------------------------------- # + +# Sentinel byte that cannot occur in normal generated template text. +_LIVE_SENTINEL = "\x00" + +# A JSON key is a double-quoted string immediately followed (after optional +# whitespace) by a colon. We only need to detect a sentinel *inside* such a +# string, so match a quoted run that ends in `":` and look for the sentinel. +_JSON_KEY_RE = re.compile(r'"((?:[^"\\]|\\.)*)"\s*:', re.S) + +_KEY_JINJA_DELIMS = ("{{", "}}", "{%", "%}", "{#", "#}") + + +def _contains_jinja_delim(text: str) -> bool: + """True if *text* contains any Jinja delimiter (live or escaped-literal).""" + return any(d in text for d in _KEY_JINJA_DELIMS) + + +def verify_no_live_jinja_in_json_keys(template_text: str) -> None: + """Reject a JSON template that carries Jinja markup in an object key. + + JinjaTurtle never templates a JSON *object key*: keys come straight from the + source and a key is an identifier/string, never a value placeholder. Any + Jinja in key position therefore means attacker-influenced source key text + reached the template. This gate fails closed on it, independent of whether a + handler left the markup *live* (a raw ``{{ ... }}`` in the key) or *escaped* + it into a ``{% raw %}`` wrapper -- both indicate a key that should never have + contained templating, so generation aborts rather than emitting it. + + Detection is done on the lexer token stream: we reconstruct the document with + every live tag collapsed to a sentinel and every ``{% raw %}``-wrapped region + also marked, then reject a sentinel that lands inside a JSON key string. + """ + import jinja2 + + env = jinja2.Environment(autoescape=True) + try: + tokens = list(env.lex(template_text)) + except jinja2.TemplateSyntaxError as exc: + raise TemplateSafetyError( + f"generated JSON template does not lex as the JinjaTurtle subset: {exc}" + ) from exc + + out: list[str] = [] + in_tag = False + for _lineno, tok_type, value in tokens: + if tok_type in ("variable_begin", "block_begin", "comment_begin"): + # A live construct: collapse to a sentinel so it is detectable if it + # sits in key position. (raw_begin/raw_end are *not* live; the data + # inside a raw block is preserved verbatim below, so an escaped key + # still shows its literal Jinja delimiters to the key check.) + in_tag = True + out.append(_LIVE_SENTINEL) + elif tok_type in ("variable_end", "block_end", "comment_end"): + in_tag = False + elif not in_tag: + # data, whitespace, raw_begin/raw_end markers, and inert raw content. + out.append(value if isinstance(value, str) else str(value)) + + reconstructed = "".join(out) + + for match in _JSON_KEY_RE.finditer(reconstructed): + key_text = match.group(1) + if _LIVE_SENTINEL in key_text or _contains_jinja_delim(key_text): + raise TemplateSafetyError( + "refusing to emit JSON template: Jinja markup appears inside a " + "JSON object key. JinjaTurtle never templates keys, so this " + "indicates attacker-influenced source key text (possible " + "template injection)." + ) + + +# +# JinjaTurtle's ERB output is produced by translating the (already-verified) +# Jinja2 subset, so the Jinja2 gate is the primary guarantee. As an independent +# ERB-side backstop we confirm that every ERB tag body is one the translator +# emits, and that no Jinja2 delimiters survived into the ERB output (which would +# indicate a raw block the translator failed to recognise -- the historical +# ``{%+ raw %}`` blind spot). +# --------------------------------------------------------------------------- # + +_ERB_TAG_RE = re.compile(r"<%[-=#]?(.*?)[-]?%>", re.S) +_JINJA_DELIMS = ("{{", "}}", "{%", "%}", "{#", "#}") + +# Bodies the ErbTranslator emits. Kept permissive for Ruby method chains it +# constructs (``@var``, ``.each_with_index``, ``JSON.generate(...)`` etc.) but +# anchored so arbitrary attacker text cannot masquerade as one. +_ERB_STMT_PATTERNS = tuple( + re.compile(p) + for p in ( + r"^require 'json'$", + r"^end$", + r"^else$", + r"^@?[A-Za-z_][\w@\.\[\]'\"]*\.each_with_index do \|[A-Za-z_]\w*, __jt_idx_\d+\| $", + r"^if .+$", + r"^elsif .+$", + r"^unless .+\.nil\?$", + r"^# Unsupported JinjaTurtle statement: .*$", + ) +) + + +def verify_erb_template_safe(template_text: str) -> None: + """Validate that *template_text* contains no leftover Jinja2 delimiters. + + The translator is the security-relevant step for ERB; this backstop ensures + no live Jinja construct survived translation (which would mean a raw block + was not recognised and source text passed through untouched). + """ + for delim in _JINJA_DELIMS: + if delim in template_text: + raise TemplateSafetyError( + "refusing to emit ERB template: it still contains the Jinja2 " + f"delimiter {delim!r}, which means source text was not fully " + "translated/neutralised (possible template injection)." + ) diff --git a/tests/test_deep_nesting_dos.py b/tests/test_deep_nesting_dos.py deleted file mode 100644 index 0ff976d..0000000 --- a/tests/test_deep_nesting_dos.py +++ /dev/null @@ -1,122 +0,0 @@ -"""Regression tests for the deeply-nested-input denial-of-service fix. - -Pathologically nested configuration (thousands of container levels) used to -drive the recursive walkers in ``core`` -- or the underlying parser itself -- -past Python's stack limit, raising an uncaught ``RecursionError``. The fix -bounds nesting depth and converts both cases into a clean, catchable -``ConfigTooDeeplyNestedError``. These tests pin that behaviour and guard -against regressions, while confirming normally-nested configs still process. -""" - -from __future__ import annotations - -import tempfile -from pathlib import Path - -import pytest - -from jinjaturtle.core import ( - ConfigTooDeeplyNestedError, - MAX_CONFIG_DEPTH, - analyze_loops, - flatten_config, - generate_ansible_yaml, - generate_jinja2_template, - parse_config, -) - - -def _run_pipeline(content: str, suffix: str) -> None: - with tempfile.NamedTemporaryFile( - "w", suffix=suffix, delete=False, encoding="utf-8" - ) as tf: - tf.write(content) - path = Path(tf.name) - try: - fmt, parsed = parse_config(path) - loops = analyze_loops(fmt, parsed) - flat = flatten_config(fmt, parsed, loops) - generate_jinja2_template( - fmt, parsed, "role", original_text=content, loop_candidates=loops - ) - generate_ansible_yaml("role", flat, loops) - finally: - path.unlink() - - -def _deep_json_objects(depth: int) -> str: - s = "0" - for _ in range(depth): - s = '{"a":' + s + "}" - return s - - -def _deep_json_arrays(depth: int) -> str: - return "[" * depth + "1" + "]" * depth - - -def _deep_xml(depth: int) -> str: - s = "v" - for _ in range(depth): - s = f"{s}" - return s - - -def _deep_yaml(depth: int) -> str: - s = "v" - for _ in range(depth): - s = "{a: " + s + "}" - return s - - -def _deep_toml(depth: int) -> str: - return "a = " + "[" * depth + "1" + "]" * depth - - -@pytest.mark.parametrize( - "content, suffix", - [ - (_deep_json_objects(5000), ".json"), - (_deep_json_arrays(100000), ".json"), - (_deep_xml(5000), ".xml"), - (_deep_yaml(3000), ".yaml"), - (_deep_toml(2000), ".toml"), - ], -) -def test_deeply_nested_input_is_rejected_cleanly(content: str, suffix: str) -> None: - # Must raise our typed error, and crucially must NOT raise RecursionError - # (pytest would report that as an error, but be explicit about intent). - with pytest.raises(ConfigTooDeeplyNestedError): - _run_pipeline(content, suffix) - - -def test_recursion_error_never_escapes() -> None: - # Belt-and-suspenders: ensure a RecursionError is never what surfaces. - content = _deep_yaml(6000) - try: - _run_pipeline(content, ".yaml") - except ConfigTooDeeplyNestedError: - pass - except RecursionError: # pragma: no cover - this is the bug we fixed - pytest.fail("RecursionError escaped instead of ConfigTooDeeplyNestedError") - - -@pytest.mark.parametrize( - "content, suffix", - [ - ('{"name":"web","port":8080,"tags":["a","b"]}', ".json"), - ("db5432", ".xml"), - ("name: web\nport: 8080\nnested:\n a:\n b: 1\n", ".yaml"), - ('[section]\nkey = "value"\nnums = [1, 2, 3]\n', ".toml"), - ("[section]\nkey = value\n", ".ini"), - ], -) -def test_normal_configs_still_process(content: str, suffix: str) -> None: - # No exception expected for ordinary, shallow configurations. - _run_pipeline(content, suffix) - - -def test_just_under_limit_is_accepted() -> None: - # A structure comfortably under the cap must still process successfully. - content = _deep_json_objects(MAX_CONFIG_DEPTH - 10) - _run_pipeline(content, ".json") diff --git a/tests/test_injection_security.py b/tests/test_injection_security.py index 0b0ee7b..4afb7a5 100644 --- a/tests/test_injection_security.py +++ b/tests/test_injection_security.py @@ -2,9 +2,9 @@ JinjaTurtle copies parts of the source config (comments, unrecognised lines, structural keys) verbatim into the generated template. If that text contains -Jinja2 delimiters it must be neutralised, otherwise attacker-influenced -config content becomes live template code when a downstream tool renders the -template. +Jinja2/ERB delimiters it must be neutralised, otherwise attacker-influenced +config content becomes live template code that executes when Ansible later +renders the template. These tests render the *generated* template the way a downstream tool would and assert that an injected payload never executes. A tripwire object is exposed @@ -24,12 +24,26 @@ import yaml as pyyaml from jinjaturtle.escape import ( escape_jinja_literal, - escape_literal, ) TRIP = "__TRIPWIRE_FIRED__" +class _UnsafeAwareLoader(pyyaml.SafeLoader): + pass + + +def _construct_unsafe(loader: _UnsafeAwareLoader, node: pyyaml.Node): + return loader.construct_scalar(node) + + +_UnsafeAwareLoader.add_constructor("!unsafe", _construct_unsafe) + + +def _safe_load_defaults(text: str): + return pyyaml.load(text, Loader=_UnsafeAwareLoader) + + class _Boom: """Returns the tripwire sentinel for any access/call an SSTI payload makes.""" @@ -83,7 +97,7 @@ def _run_jinjaturtle(tmp_path: Path, source_name: str, body: str, fmt: str): text=True, ) assert res.returncode == 0, f"generation failed: {res.stderr}" - defaults = pyyaml.safe_load(dfl.read_text()) or {} + defaults = _safe_load_defaults(dfl.read_text()) or {} return tpl.read_text(), defaults @@ -153,6 +167,86 @@ def test_generated_template_is_renderable(tmp_path, fmt, name, body): _render_jinja(template_text, defaults) +# --- JSON object-key injection ------------------------------------------------ +# +# The JSON handler copies the text *between* scalar values (object keys, +# punctuation) verbatim. A key is never a value placeholder, so Jinja markup in a +# key can only come from attacker-influenced source text. The output gate +# (verify_no_live_jinja_in_json_keys) must fail closed on it -- including the +# "benign-looking name" form (e.g. ``{{ ansible_hostname }}``) that the generic +# allowlist would otherwise accept as an ordinary variable reference, and which +# could leak an in-scope variable's value into the rendered config at apply time. + +JSON_KEY_INJECTION_BODIES = [ + # benign-looking variable reference (the residual bypass: passes the generic + # allowlist but must still be rejected in *key* position) + '{ "{{ ansible_hostname }}": "v" }', + # dotted reference (e.g. dumping another host's vars) + '{ "{{ hostvars.localhost }}": 1 }', + # self-referencing a sibling-derived variable name + '{ "{{ role_port }}": "x", "port": 8080 }', + # classic gadget (already rejected historically; kept as a guard) + '{ "{{ cycler.__init__.__globals__ }}": 1 }', + # statement injection in a key + '{ "{% for x in y %}k{% endfor %}": 1 }', + # nested object key + '{ "ok": { "{{ ansible_hostname }}": 2 } }', +] + + +@pytest.mark.parametrize("body", JSON_KEY_INJECTION_BODIES) +def test_json_key_injection_fails_closed_cli(tmp_path, body): + src = tmp_path / "evil.json" + src.write_text(body, encoding="utf-8") + out = tmp_path / "out.j2" + res = subprocess.run( + [ + sys.executable, + "-m", + "jinjaturtle.cli", + str(src), + "-f", + "json", + "--role-name", + "role", + "-t", + str(out), + ], + capture_output=True, + text=True, + ) + assert res.returncode == 2, f"expected fail-closed, got rc={res.returncode}" + assert "refusing to generate unsafe template" in res.stderr + assert not out.exists(), "no template may be written when the gate refuses" + + +def test_json_benign_keys_still_generate(tmp_path): + src = tmp_path / "ok.json" + src.write_text('{ "host": "localhost", "port": 8080 }', encoding="utf-8") + out = tmp_path / "out.j2" + res = subprocess.run( + [ + sys.executable, + "-m", + "jinjaturtle.cli", + str(src), + "-f", + "json", + "--role-name", + "demo", + "-t", + str(out), + ], + capture_output=True, + text=True, + ) + assert res.returncode == 0, res.stderr + template_text = out.read_text() + # Keys stay literal; values become placeholders. + assert '"host":' in template_text + assert "demo_host" in template_text + + # --- Unit-level guarantees for the escaper itself --------------------------- SSTI_PAYLOADS = [ @@ -231,17 +325,3 @@ def test_endraw_whitespace_control_cannot_break_out(endraw): rendered = env.from_string(escaped).render(boom=_Boom()) assert TRIP not in rendered assert rendered == payload - - -def test_escape_literal_is_jinja_alias(): - """The public ``escape_literal`` wrapper is Jinja2-only.""" - payload = "{{ 7*7 }}" - assert escape_literal(payload) == escape_jinja_literal(payload) - - -def test_escape_literal_jinja_output_is_inert(): - """End-to-end: text routed through escape_literal does not execute.""" - escaped = escape_literal("{{ boom.run('x') }}") - env = jinja2.Environment(undefined=jinja2.ChainableUndefined) - rendered = env.from_string(escaped).render(boom=_Boom()) - assert TRIP not in rendered diff --git a/tests/test_output_safety_gate.py b/tests/test_output_safety_gate.py index 813ded0..055ceb3 100644 --- a/tests/test_output_safety_gate.py +++ b/tests/test_output_safety_gate.py @@ -25,6 +25,7 @@ from jinjaturtle.multi import process_directory from jinjaturtle.safety import ( TemplateSafetyError, verify_jinja2_template_safe, + verify_erb_template_safe, ) @@ -179,6 +180,20 @@ def _strip_raw_blocks(text: str) -> str: ) +# --------------------------------------------------------------------------- # +# ERB gate. +# --------------------------------------------------------------------------- # + + +def test_erb_gate_blocks_leftover_jinja_delimiters(): + with pytest.raises(TemplateSafetyError): + verify_erb_template_safe("ok <%= @x %> but {{ leftover }} here") + + +def test_erb_gate_allows_clean_erb(): + verify_erb_template_safe("memory = <%= @memory_limit %>\n") + + # --------------------------------------------------------------------------- # # End-to-end CLI: fail closed. # --------------------------------------------------------------------------- # diff --git a/tests/test_release_blockers.py b/tests/test_release_blockers.py deleted file mode 100644 index c4d3a63..0000000 --- a/tests/test_release_blockers.py +++ /dev/null @@ -1,98 +0,0 @@ -from pathlib import Path - -import pytest -import yaml -from defusedxml.common import DTDForbidden - -from jinjaturtle.core import ( - VariableNameCollisionError, - generate_ansible_yaml, - parse_config, -) -from jinjaturtle.handlers.xml import XmlHandler -from jinjaturtle.multi import iter_supported_files, process_directory - - -def test_defaults_string_values_are_ansible_unsafe(): - output = generate_ansible_yaml( - "role", [(("section", "key"), "{{ lookup('pipe', 'id') }}")] - ) - assert "!unsafe" in output - assert yaml.safe_load(output)["role_section_key"] == "{{ lookup('pipe', 'id') }}" - - -def test_variable_name_collisions_fail_closed(): - with pytest.raises(VariableNameCollisionError): - generate_ansible_yaml( - "role", - [ - (("section", "a-b"), "safe"), - (("section", "a_b"), "evil"), - ], - ) - - -def test_folder_mode_does_not_follow_file_symlinks(tmp_path: Path): - outside = tmp_path / "outside.ini" - outside.write_text("[s]\nsecret = value\n", encoding="utf-8") - root = tmp_path / "root" - root.mkdir() - (root / "link.ini").symlink_to(outside) - (root / "real.ini").write_text("[s]\nname = ok\n", encoding="utf-8") - - files = iter_supported_files(root, recursive=True) - assert files == [root / "real.ini"] - defaults, _outputs = process_directory(root, recursive=True, role_prefix="role") - assert "secret" not in defaults - assert "real.ini" in defaults - - -def test_xml_parser_rejects_dtd_and_entities(tmp_path: Path): - xml_path = tmp_path / "evil.xml" - xml_path.write_text(']>&x;', encoding="utf-8") - - with pytest.raises(DTDForbidden): - parse_config(xml_path) - - -def test_xml_source_comments_cannot_forge_internal_markers(): - source = ( - "ok" - ) - template = XmlHandler().generate_jinja2_template(None, "role", original_text=source) - - assert "{% if role_missing is defined %}" not in template - assert "{% endif %}" not in template - assert "IF:role_missing" in template - assert "ENDIF:role_missing" in template - - -def test_cli_refuses_to_write_output_symlink(tmp_path: Path): - import subprocess - import sys - - src = tmp_path / "input.ini" - src.write_text("[s]\nname = ok\n", encoding="utf-8") - target = tmp_path / "target.yml" - target.write_text("do not overwrite\n", encoding="utf-8") - link = tmp_path / "defaults.yml" - link.symlink_to(target) - - result = subprocess.run( - [ - sys.executable, - "-m", - "jinjaturtle.cli", - str(src), - "-d", - str(link), - ], - text=True, - stdout=subprocess.PIPE, - stderr=subprocess.PIPE, - env={"PYTHONPATH": str(Path(__file__).resolve().parents[1] / "src")}, - check=False, - ) - - assert result.returncode != 0 - assert target.read_text(encoding="utf-8") == "do not overwrite\n" diff --git a/tests/test_security_hardening.py b/tests/test_security_hardening.py new file mode 100644 index 0000000..41a6b05 --- /dev/null +++ b/tests/test_security_hardening.py @@ -0,0 +1,120 @@ +from __future__ import annotations + +from pathlib import Path +import os + +import pytest +import yaml +from defusedxml.common import EntitiesForbidden + +from jinjaturtle import cli +from jinjaturtle.core import generate_ansible_yaml, parse_config, flatten_config +from jinjaturtle.multi import is_supported_file, iter_supported_files, process_directory +from jinjaturtle.output_safety import OutputPathError, write_text_safely + + +class UnsafeAwareLoader(yaml.SafeLoader): + pass + + +def _unsafe(loader: UnsafeAwareLoader, node: yaml.Node): + return loader.construct_scalar(node) + + +UnsafeAwareLoader.add_constructor("!unsafe", _unsafe) + + +def test_jinja_values_are_emitted_as_ansible_unsafe(tmp_path: Path): + src = tmp_path / "app.ini" + src.write_text("[main]\ncmd = {{ lookup('pipe','id') }}\n", encoding="utf-8") + + fmt, parsed = parse_config(src) + defaults_yaml = generate_ansible_yaml("role", flatten_config(fmt, parsed)) + + assert "role_main_cmd: !unsafe" in defaults_yaml + assert "{{ lookup(''pipe'',''id'') }}" in defaults_yaml + loaded = yaml.load(defaults_yaml, Loader=UnsafeAwareLoader) + assert loaded["role_main_cmd"] == "{{ lookup('pipe','id') }}" + + +def test_folder_mode_marks_nested_jinja_values_and_ids_unsafe(tmp_path: Path): + src = tmp_path / "src" + src.mkdir() + # Filename ids are source-derived values too. + (src / "{{ bad }}.yaml").write_text( + "message: \"{{ lookup('pipe','id') }}\"\n", encoding="utf-8" + ) + + defaults_yaml, _outputs = process_directory( + src, recursive=False, role_prefix="role" + ) + + assert "id: !unsafe" in defaults_yaml + assert "role_message: !unsafe" in defaults_yaml + loaded = yaml.load(defaults_yaml, Loader=UnsafeAwareLoader) + assert loaded["role_items"][0]["id"] == "{{ bad }}.yaml" + assert loaded["role_items"][0]["role_message"] == "{{ lookup('pipe','id') }}" + + +def test_xml_parser_rejects_entities_when_called_as_library(tmp_path: Path): + src = tmp_path / "bad.xml" + src.write_text( + "]>&xxe;", + encoding="utf-8", + ) + + with pytest.raises(EntitiesForbidden): + parse_config(src, "xml") + + +def test_folder_mode_does_not_follow_symlinked_files(tmp_path: Path): + real = tmp_path / "secret.ini" + real.write_text("[main]\nsecret=yes\n", encoding="utf-8") + root = tmp_path / "root" + root.mkdir() + link = root / "link.ini" + link.symlink_to(real) + + assert not is_supported_file(link) + assert iter_supported_files(root, recursive=False) == [] + assert iter_supported_files(root, recursive=True) == [] + + +@pytest.mark.skipif(not hasattr(os, "symlink"), reason="symlinks unavailable") +def test_cli_refuses_to_write_through_final_symlink(tmp_path: Path): + target = tmp_path / "target.txt" + target.write_text("keep\n", encoding="utf-8") + link = tmp_path / "out.yml" + link.symlink_to(target) + + with pytest.raises(OutputPathError): + write_text_safely(link, "replace\n") + + assert target.read_text(encoding="utf-8") == "keep\n" + assert link.is_symlink() + + +def test_cli_refuses_symlinked_output_parent(tmp_path: Path): + real_dir = tmp_path / "real" + real_dir.mkdir() + link_dir = tmp_path / "linkdir" + link_dir.symlink_to(real_dir, target_is_directory=True) + + with pytest.raises(OutputPathError): + write_text_safely(link_dir / "out.yml", "data\n") + + assert not (real_dir / "out.yml").exists() + + +def test_cli_reports_unsafe_output_path_without_overwriting_symlink(tmp_path: Path): + cfg = tmp_path / "app.ini" + cfg.write_text("[main]\nname = ok\n", encoding="utf-8") + target = tmp_path / "target.yml" + target.write_text("keep\n", encoding="utf-8") + link = tmp_path / "defaults.yml" + link.symlink_to(target) + + exit_code = cli._main([str(cfg), "--defaults-output", str(link)]) + + assert exit_code == 2 + assert target.read_text(encoding="utf-8") == "keep\n" diff --git a/tests/test_yaml_handler.py b/tests/test_yaml_handler.py index 543c759..548f5a6 100644 --- a/tests/test_yaml_handler.py +++ b/tests/test_yaml_handler.py @@ -177,15 +177,17 @@ def test_yaml_loop_preserves_blank_separator_after_list(tmp_path: Path): "blank", "require:\n - rubocop-performance\n - rubocop-rspec\n\nAllCops:\n NewCops: enable\n", "\n - rubocop-rspec\n\nAllCops:", + "<% end %>\nAllCops:", ), ( "no_blank", "require:\n - rubocop-performance\n - rubocop-rspec\nAllCops:\n NewCops: enable\n", "\n - rubocop-rspec\nAllCops:", + "<% end %>AllCops:", ), ] - for label, text, rendered_expected in cases: + for label, text, rendered_expected, erb_expected in cases: path = tmp_path / f"{label}.yml" path.write_text(text, encoding="utf-8")