Posts Tagged ‘Python’
[PyDataGlobal2025] Tools, Empathy, and the Craft of Building Delightful Data Experiences
Lecturer
Isabel Zimmerman is a Senior Software Engineer at Posit, PBC (formerly RStudio). She was the first full-time Python open-source hire at the company and began her tenure building MLOps packages before shifting focus to the Python experience inside interactive development environments. Her current work centers on Positron, a next-generation data-science IDE. Beyond computing she is an avid fantasy reader and bookbinder, interests that inform her view of tools as objects that can carry quiet power across generations of users.
Abstract
Every practitioner occupies a position on the continuum between tool user and tool builder. This keynote explores that continuum through the dual lenses of technical excellence and human empathy. Drawing on concrete examples from the Positron IDE and the broader open-source Python ecosystem, it articulates a set of “hard skills” (modularity, reproducibility, flexibility) and “soft skills” (knowing the user, discoverability, small improvements with large impact, and explaining one’s work). The argument is that tools become delightful only when both categories are deliberately cultivated, and that the barrier to becoming a builder has never been lower.
From Consumer to Creator: Reframing Everyday Practice
A tool is defined simply as anything that carries out a particular function. Under that definition most data scientists already build tools—whether a Git alias that corrects a habitual typo, a reusable function shared in Slack, a dashboard that informs business decisions, or a private utility that solves a personal measurement problem. The psychological barrier that prevents many practitioners from identifying as builders is therefore largely artificial. Framing the act of extraction and encapsulation as tool construction lowers that barrier and simultaneously improves personal productivity and future reproducibility.
The transition from pure consumer to occasional creator is further eased by contemporary language models. Functions that once required manual packaging can now be sketched in natural language and refined iteratively. The resulting artifacts need not be public; a private package that accelerates one’s own daily workflow is already a contribution to the wider ecosystem because it reduces friction for at least one user—oneself.
Hard Skills of Tool Design
Three technical properties form the backbone of robust tools. Modularity allows a system to grow with its users. By leaning on existing community infrastructure—FastAPI for REST endpoints, Code OSS for the editor substrate—builders can concentrate effort on the distinctive value they wish to add. The same modular surface also supplies clear extension points, encouraging specialized packages that solve narrow, high-value problems.
Reproducibility remains a foundational requirement of trustworthy science. Graphical exploration interfaces are powerful, yet they risk introducing non-reproducible click sequences. Positron’s data explorer illustrates one resolution: every filter and sort operation is internally represented so that a single button can emit executable code that recreates the identical view. The cycle of exploration is thereby closed inside a language rather than left as a sequence of manual steps.
Flexibility must be tempered by the Zen of Python’s preference for simplicity. Functions that accept an ever-expanding union of input types quickly become unmaintainable. Preferring a small number of well-defined entry points and composing them later yields systems that remain extensible without collapsing under their own complexity. Context windows supplied to language models follow the same principle: start with a carefully chosen default set of information and allow the user to add or remove context explicitly.
Soft Skills and the Human Side of Interfaces
Technical excellence alone does not produce tools that people love. Empathy for the intended user is equally decisive. Data work is characterized by iterative exploration of uncharted territory, whereas classical software engineering often constructs well-specified structures in known domains. An interface optimized solely for the latter will frustrate the former. Permanent, always-available consoles, column-aware completions, and language-server optimizations tuned to data-frame idioms are concrete expressions of that empathy.
Discoverability ensures that high-impact features do not remain secret passages. Action bars that surface “render on save,” one-click code-cell insertion, and help panes that render richly formatted docstrings bring frequently needed capabilities into immediate view. Small ergonomic improvements—running a Streamlit or Dash application with the correct launcher rather than a plain Python invocation—accumulate into large reductions in daily friction.
Finally, the act of explaining one’s work closes a vital feedback loop. Writing documentation, type annotations, or even lightweight notes in a project file forces clarity of thought. The same artifacts later serve both future collaborators and future selves. The principle “if your writing helps even one person it is worth doing, especially if that person is you” applies equally to private architectural notes and public getting-started guides.
Closing the Loop Between Building and Using
Tools improve through continuous cycles of use, observation of pain points, and iterative refinement. Feedback—whether GitHub issues, hallway conversations, or structured user testing—supplies the raw material for those cycles. Because every practitioner is simultaneously a consumer and a potential contributor, each unique perspective enriches the shared ecosystem. The mission is not the construction of a final, perfect package but the ongoing cultivation of experiences that feel beautiful, empowering, and precisely fitted to the work at hand.
Links:
[VoxxedDaysLuxemburg2026] Introduction to Machine Learning for Software Engineers: A Comprehensive Framework from Data Pre-processing to Responsible Deployment
Lecturer
G. Darwish is a software engineer operating within Lunat in the Netherlands. Holding a Master’s degree in Artificial Intelligence, his specialized technical focus lies in trustworthy AI frameworks, predictive modeling, and the evolving regulatory landscape surrounding European Union AI policy. Beyond practical software development, his work addresses algorithmic accountability, mitigation of model bias, and the operational deployment of supervised learning systems within enterprise environments.
Abstract
This paper presents a rigorous, end-to-end framework for integrating traditional supervised machine learning methodologies into modern software engineering workflows. Moving beyond high-level artificial intelligence discourse, it details the mathematical and operational distinctions between classical deterministic programming and empirical pattern learning. Utilizing the canonical 1994 UCI Adult Income dataset as a case study, the investigation explores exploratory data analysis (EDA), data cleaning, categorical encoding, feature scaling, and feature engineering. It addresses the trade-offs inherent in model selection, regularization, and hyperparameter optimization to balance accuracy against explainability. Furthermore, the study formalizes performance evaluation through confusion matrices, precision, recall, and F1-scores, while confronting the sociotechnical challenge of algorithmic bias. Finally, it outlines industrial deployment protocols, focusing on CI/CD release gates, data drift detection, and continuous monitoring paradigms necessary for maintaining robust, trustworthy machine learning systems in production.
Technical Context: Paradigm Shift from Deterministic Software to Empirical Learning
Traditional software engineering relies on deterministic paradigms where explicit, domain-specific rules are authored by engineers. Input data is processed through these predefined rules to yield deterministic outputs. However, complex real-world tasks—such as visual object recognition, natural language comprehension, and dynamic fraud detection—present rule sets of such high dimensionality and edge-case density that explicit manual programming becomes intractable.
+---------------------------------------------+
| Traditional Programming |
| Input Data + Explicit Rules ---> Output |
+---------------------------------------------+
| Machine Learning |
| Input Data + Output ---> Learned Rules |
+---------------------------------------------+
Machine learning reorganizes this computational paradigm. Rather than manually codifying decision logic, supervised learning algorithms consume historical inputs alongside validated outputs (ground truth labels) to synthesize an internal numerical representation of the underlying patterns.
# Deterministic Rule-Based Paradigm
def evaluate_loan_application(income, score):
if income > 50000 and score > 700:
return "APPROVED"
return "REJECTED"
# Empirical Machine Learning Paradigm
from sklearn.linear_model import LogisticRegression
def train_ml_classifier(X_train, y_train):
model = LogisticRegression(C=1.0)
model.fit(X_train, y_train)
return model
To maintain technical precision, software architectures must distinguish between functional tiers within the artificial intelligence ecosystem:
- Artificial Intelligence (AI): The broad domain encompassing any artificial system capable of exhibiting task intelligence, spanning rule engines, heuristic search solvers, and statistical estimators.
- Narrow AI versus General AI (AGI): Narrow AI designates systems engineered and optimized to execute a singular, highly scoped task (such as credit evaluation or image classification). Artificial General Intelligence (AGI) implies systems possessing domain-agnostic conceptualization and autonomous reasoning across disparate cognitive spaces.
- Machine Learning (ML): A subdiscipline of AI focused on algorithms that optimize performance parameters through statistical exposure to empirical data.
- Deep Learning & Generative AI: Specialized subsets of ML utilizing multi-layered neural networks (e.g., Transformer architectures) capable of hierarchical abstraction and synthesis of novel text, image, or structural artifacts.
Exploratory Data Analysis and Pipeline Engineering
Data preparation constitutes the primary deterministic driver of machine learning performance. Model optimization relies entirely on the structural integrity of the input data. The primary domain of reference analyzed throughout this pipeline is the UCI Adult Income dataset, containing structural socio-demographic features designed to predict whether an individual’s annual income exceeds $50,000.
+---------------------------------------------+
| Machine Learning Pipeline |
| |
| [ Ingest Data ] |
| | |
| v |
| [ EDA & Data Prep ] |
| | |
| v |
| [ Categorical Encoding ] |
| | |
| v |
| [ Feature Scaling ] |
| | |
| v |
| [ Model Training & Evaluation ] |
| | |
| v |
| [ Deployment & Monitoring ] |
+---------------------------------------------+
Data Cleansing and Imputation
Raw datasets frequently exhibit missing entries, structural anomalies, and non-conforming placeholder values. In complete feature sets, missing indices marked by symbols such as question marks must be converted to native null types. Engineers must decide between two primary mitigation paths:
- Row Excision: Removing observations containing null values when the missing subset constitutes a minor percentage of the total dataset, thereby preserving feature distribution without introducing artificial bias.
- Statistical Imputation: Substituting missing attributes with central tendency metrics (mean, median, or mode) or inferring values via auxiliary regression models when data volume retention is critical.
import pandas as pd
import numpy as np
# Ingestion and clean-up of sentinel values
df = pd.read_csv("adult_income.csv")
df.replace("?", np.nan, inplace=True)
df.dropna(inplace=True)
# Target vector binary mapping
df["target"] = (df["income"] == ">50K").astype(int)
Feature Encoding Techniques
Algorithms process numerical vectors; therefore, qualitative textual fields must undergo rigorous mathematical transformation.
- One-Hot Encoding: Applied to low-cardinality nominal variables (such as education status or relationship type). This operation converts a categorical feature containing N distinct values into N distinct binary vector columns containing mutually exclusive 0 or 1 indicators.
- High-Cardinality Scaling: Applied when categorical features possess dozens or hundreds of unique entries (e.g., native country). Here, frequency encoding or target encoding is utilized to project categories into a bounded numeric spectrum between 0 and 1, mitigating dimensional explosion.
# One-Hot Encoding implementation
encoded_df = pd.get_dummies(
df,
columns=["education", "workclass"],
drop_first=True
)
Feature Scaling and Vector Normalization
When numerical features possess wildly disparate ranges—such as age (17 to 90) versus weekly work hours (1 to 99) or capital gains (0 to 99,999)—gradient-based optimization algorithms suffer from unstable weight updates. Models over-index on raw magnitude rather than structural correlation.
- Min-Max Scaling: Rescales values linearly to force the feature domain strictly within [0, 1]:
X_norm = (X - X_min) / (X_max - X_min) - Standardization (Z-Score Normalization): Centers data around a zero mean with unit variance, robustifying the system against outliers:
X_std = (X - mean) / standard_deviation
Feature Engineering
Engineers extract amplified signals by composing derived variables from underlying raw dimensions. For instance, raw continuous metrics like weekly working hours can be binned into discretized operational states (such as part-time, standard, or overtime). Similarly, capital gains and capital losses can be integrated into a unified boolean feature tracking net capital activity.
# Constructing explicit engineered signals
df["capital_active"] = (
(df["capital_gain"] > 0) |
(df["capital_loss"] > 0)
).astype(int)
df["overtime_worker"] = (
df["hours_per_week"] > 40
).astype(int)
Empirical Model Architecture, Generalization, and Optimization
Generalization, Overfitting, and Underfitting
The core objective of machine learning engineering is to build models that demonstrate high generalization performance on unseen production data. High accuracy on training data is uninformative if the underlying functional representation fails under novel conditions.
Underfitting (High Bias)
+---------------------------------------------+
| o o |
| \ |
| \----o |
| \---o |
+---------------------------------------------+
Simplistic fit fails true trend
Balanced Generalization
+---------------------------------------------+
| o / o |
| \ / |
| \---o |
| \---o |
+---------------------------------------------+
Captures underlying structural trend
Overfitting (High Variance)
+---------------------------------------------+
| o----\ /--o |
| \-/ |
| o------------------o----o |
+---------------------------------------------+
Fits noise and fails to generalize
- Underfitting (High Bias): Occurs when the decision boundary is excessively simplistic (e.g., fitting a linear model to non-linear parabolic data), preventing the algorithm from capturing fundamental data relationships.
- Overfitting (High Variance): Occurs when a hyper-complex decision boundary memorizes noisy anomalies and fine-grained variations specific to the training set. While training performance reaches optimal metrics, validation accuracy drops significantly when evaluated against new inputs.
To preserve operational generalization, training strategies require splitting the raw dataset into three distinct partitions: an 80% Training Set (to optimize internal parameters), a 10% Validation Set (to iterate on hyperparameters), and a 10% Test Set (held back to measure generalized accuracy prior to release). Stratification must be maintained across splits to mirror real-world label distributions.
from sklearn.model_selection import train_test_split
X = encoded_df.drop(columns=["target", "income"])
y = encoded_df["target"]
# Stratified multi-tier data partitioning
X_train, X_temp, y_train, y_temp = (
train_test_split(
X, y,
test_size=0.2,
stratify=y,
random_state=42
)
)
X_val, X_test, y_val, y_test = (
train_test_split(
X_temp, y_temp,
test_size=0.5,
stratify=y_temp,
random_state=42
)
)
Architectural Classification Algorithms
Selection of mathematical architectures depends on explicit problem constraints, interpretability bounds, and data volume:
- Linear Regression: Maps independent variables linearly to continuous targets (
y = a*x + b), serving as a baseline for numerical estimation. - Logistic Regression: Applies a sigmoid activation function over a linear combination of inputs, squeezing continuous outputs into a probability spectrum between 0 and 1 to establish binary classification thresholds.
- Decision Trees: Sequentially partitions feature spaces using calculated entropy reduction or Gini impurity thresholds. Highly interpretable as nested conditional logic, but susceptible to severe overfitting if left unpruned.
- K-Nearest Neighbors (KNN): A non-parametric instance-based classifier that maps new inputs to the majority label among its K nearest geometric neighbors within vector space. Computationally expensive during inference on large datasets.
from sklearn.neighbors import KNeighborsClassifier
from sklearn.tree import DecisionTreeClassifier
# Baseline classification architectures
logistic_clf = LogisticRegression(max_iter=1000)
tree_clf = DecisionTreeClassifier(max_depth=5)
knn_clf = KNeighborsClassifier(n_neighbors=5)
Regularization and Hyperparameter Search
Regularization injects explicit loss penalties to constrain model complexity. L1 Regularization (Lasso) shrinks irrelevant feature weights strictly to zero, effectively performing automatic feature selection. L2 Regularization (Ridge) penalizes large squared weight magnitudes, distributing importance evenly across features to prevent individual variables from dominating decision boundaries.
Hyperparameters—such as decision tree depth bounds or KNN neighborhood sizes (K)—cannot be learned directly via gradient descent. Engineers deploy systematically structured parameter searches (e.g., Grid Search Cross-Validation) across validation sets to isolate optimal configurations.
from sklearn.model_selection import GridSearchCV
# Systematic Hyperparameter Search
param_grid = {
'C': [0.01, 0.1, 1.0, 10.0],
'penalty': ['l2']
}
grid_search = GridSearchCV(
estimator=LogisticRegression(max_iter=1000),
param_grid=param_grid,
cv=5,
scoring='f1'
)
grid_search.fit(X_train, y_train)
best_model = grid_search.best_estimator_
Evaluation Frameworks and Decision-Making Diagnostics
Evaluation based solely on raw accuracy is fundamentally misleading when dealing with imbalanced datasets. If an income dataset contains 74% low-earning records, a trivial dummy model that predicts “low income” across all inputs achieves an artificial 74% accuracy while lacking true predictive capability.
+---------------------------------------------+
| ACTUAL CLASS |
| Pos (>50K) | Neg (<=50K) |
+-----------------------+---------------------+
| PREDICTED Positive | PREDICTED Negative |
| True Pos (TP) | False Neg (FN) |
| False Pos (FP) | True Neg (TN) |
+-----------------------+---------------------+
Formal Evaluation Metrics
Detailed evaluation relies on metrics derived from the Confusion Matrix:
- Accuracy: The basic ratio of correct classifications over total evaluations:
Accuracy = (TP + TN) / (TP + TN + FP + FN) - Precision: Measures the exactness of positive classifications. High precision minimizes False Positives (crucial in spam filtering or loan approvals where misclassifying an unqualified candidate introduces financial risk):
Precision = TP / (TP + FP) - Recall (Sensitivity): Measures the ability to capture all true positive cases. High recall minimizes False Negatives (essential in cancer detection or fraud alerts where missing a positive case carries severe consequences):
Recall = TP / (TP + FN) - F1-Score: The harmonic mean balancing Precision and Recall into a single metric for comparing imbalanced models:
F1-Score = 2 * (Precision * Recall) / (Precision + Recall)
from sklearn.metrics import (
classification_report,
confusion_matrix
)
y_pred = best_model.predict(X_test)
# Display diagnostic metrics
print("Confusion Matrix:")
print(confusion_matrix(y_test, y_pred))
print("\nClassification Metrics:")
print(classification_report(y_test, y_pred))
Algorithmic Bias, Fairness Metrics, and Remediation Strategies
Machine learning models absorb, codify, and scale historical human biases embedded within training data. Discarding explicit sensitive identifiers (e.g., race, gender, or age) is insufficient to guarantee fairness. Secondary features (such as postal code or historical employment category) act as proxies, enabling algorithms to reconstruct demographic biases through latent data correlations.
+---------------------------------------------+
| Bias Mitigation Lifecycles |
| |
| 1. Pre-Processing |
| - Resampling & Weight Adjustment |
| |
| 2. In-Processing |
| - Fairness Penalties Added to Loss |
| |
| 3. Post-Processing |
| - Group-Specific Decision Bounds |
+---------------------------------------------+
Disparate Impact and Mathematical Fairness Metrics
Fairness must be systematically quantified across sensitive sub-groups:
- Demographic Parity: Requires equal selection rates across sensitive groups regardless of underlying baseline differences:
P(Predicted = 1 | Group A) = P(Predicted = 1 | Group B) - Equalized Odds: Requires equivalent error rates across groups, mandating equal True Positive Rates (TPR) and equal False Positive Rates (FPR):
P(Predicted = 1 | Actual = 1, Group A) = P(Predicted = 1 | Actual = 1, Group B)
Remediation Strategies
- Pre-Processing Mitigation: Modifies training sample distributions by re-weighting or oversampling underrepresented demographics before model fitting.
- In-Processing Mitigation: Injects structural fairness constraints directly into the objective loss function. The algorithm is explicitly penalized when optimization steps increase parity gaps between demographic groups.
- Post-Processing Mitigation: Alters decision threshold parameters independently for different demographic sub-groups post-training to satisfy target equity metrics.
# Utilizing Fairlearn for Bias Remediation
from fairlearn.reductions import (
ExponentiatedGradient,
DemographicParity
)
# Define fairness constraints
mitigated_engine = ExponentiatedGradient(
estimator=LogisticRegression(max_iter=1000),
constraints=DemographicParity()
)
# Train with sensitive features
mitigated_engine.fit(
X_train,
y_train,
sensitive_features=sensitive_train
)
Mitigating algorithmic bias introduces an operational trade-off: enforcing tighter demographic constraints can reduce aggregate accuracy scores. Product engineering teams must weigh these performance drop-offs against legal compliance standards, ethical responsibilities, and corporate deployment policies.
MLOps: Production Deployment, CI/CD Gates, and Continuous Monitoring
Moving a model from an experimental Jupyter Notebook into a reliable production architecture requires robust MLOps practices. In production, model artifacts are essentially serialized weight configurations (e.g., Pickle files or GGUF structures) that execute within wrapped microservices.
+---------------------------------------------+
| Production MLOps Pipeline |
| |
| [ Model Registry (Weights) ] |
| | |
| v |
| [ Automated CI/CD Gates ] |
| | |
| v |
| [ Inference Service Endpoint ] |
| | |
| v |
| [ Drift Dashboard & Alert Triggers ] |
+---------------------------------------------+
Automated CI/CD Release Gates
Automated continuous integration and deployment pipelines must execute rigorous validation suites before any candidate model artifact is deployed:
- Performance Thresholds: Automated checks block deployments if validation F1-scores drop below predefined baselines (e.g., F1 < 0.60).
- Fairness Audit Gates: Pipelines fail build processes if the calculated true positive rate divergence across sensitive demographic groups exceeds strict limits (e.g., Delta TPR > 0.05).
- Schema Integrity Rules: Ingestion pipelines validate incoming payloads to catch schema modifications, missing fields, or unexpected data types before hitting model boundaries.
Drift Detection and Telemetry
Once operational, production models face continuous environment degradation:
- Data Drift: Occurs when input distributions shift over time (e.g., macroeconomic fluctuations changing baseline salary levels) while the underlying target relationships remain constant.
- Concept Drift: Occurs when the fundamental statistical relationship between input features and target outputs changes entirely (e.g., consumer behavior shifts following major regulatory adjustments).
- Adversarial Poisoning: Intentionally manipulated payload streams designed to corrupt learning models or exploit decision boundaries.
Engineering teams must log inference inputs, prediction outputs, and feature distributions in continuous monitoring systems. When metrics exceed statistical drift thresholds, automated alerts trigger secondary retraining pipelines, model registry rollbacks, or fallback to deterministic logic.
Links
[DevoxxUK2026] Aspiring Speakers: Learning Python to Buy Shoes
Lecturer
Isaac Oldwood is an emerging software engineer and public speaker passionate about accessible, project-based learning. His journey from mathematics student to developer exemplifies practical skill acquisition through real-world problem-solving.
Abstract
Isaac Oldwood shares a personal narrative of mastering Python by tackling an everyday challenge: automating the purchase of limited-edition sneakers. This beginner-friendly account traces the evolution from rudimentary automation scripts to sophisticated API interactions, highlighting key lessons on learning through failure, iteration, and building relevant projects.
From MATLAB to Real-World Automation: A Developer’s Origin Story
Beginning in 2016 as a first-year mathematics student at the University of Nottingham, Isaac encountered programming through a compulsory MATLAB module. Recognizing the value of coding skills for modern mathematicians, he excelled yet questioned his academic path upon receiving strong results. Seeking guidance through traditional searches, he discovered that foundational knowledge in variables and control structures should pair with building practical projects aligned with personal interests.
The chosen project involved acquiring Yeezy Boost 350 V2 “Bred” sneakers. Initial attempts employed PyAutoGUI for keyboard and mouse automation, relying on screen coordinates and arbitrary sleep timers. These approaches proved brittle against dynamic web interfaces and competitive release timings.
Subsequent iterations leveraged Selenium for browser automation, enabling direct interaction with HTML elements and conditional waits. This advancement improved reliability as scripts adapted to page changes. Further research into inter-computer communication revealed APIs, leading to direct HTTP interactions using the Requests library. By reverse-engineering checkout flows, Isaac implemented PUT requests to cart endpoints and POST requests for payment processing, dramatically reducing latency compared to full page loads burdened by images, trackers, and fonts.
Scalability challenges emerged when extending the solution to housemates. Synchronous execution created unfair queuing. Transitioning to asynchronous operations with HTTPX allowed concurrent checkouts, ensuring equitable opportunity at release moments.
Despite technical refinements, the script ultimately failed against robust anti-bot measures deployed by the retailer. Success arrived through manual effort: queuing physically outside a store at dawn. This outcome reinforced that automation serves as a learning vehicle rather than a guaranteed solution.
Conclusion
Isaac’s engaging story underscores fundamental truths about technical education. Learning proves enjoyable when rooted in passion projects. Failure constitutes an integral component of growth, yielding deeper insights than initial successes. Building tangible solutions to personal problems accelerates skill development far beyond theoretical study. Aspiring developers benefit immensely from identifying relevant challenges and iterating relentlessly toward mastery.
Links:
[PyConUS2025] Why `len(‘😶🌫️’) == 4` and Other Unexpected Behaviors of Python Strings
Lecturer
Marie Roald is a researcher, data scientist, and educator affiliated with the Norwegian Language Bank at the National Library of Norway. She has more than eight years of experience teaching Python to secondary-school students, teachers, and professionals, and is a co-founder and organizer of PyLadies Oslo. Yngve Mardal Moe is an experienced Python educator, developer, and data-science consultant who previously led the redesign of an introductory Python course at the Norwegian University of Life Sciences; he currently serves as tech lead on automation projects for the Norwegian power grid. Together they bring complementary perspectives from language technology and software engineering to the practical difficulties of Unicode handling.
Abstract
Python strings appear simple until everyday operations—length measurement, equality testing, case conversion, and slicing—produce counter-intuitive results. This presentation traces those surprises to the underlying Unicode encoding model, the distinction between code points and grapheme clusters, the existence of multiple normalized forms, and the incomplete implementation of locale-sensitive operations in the language. The speakers supply concrete recommendations for robust comparison, normalization, and length measurement that avoid the most common pitfalls.
Encoding, Code Points, and the Limits of Naïve Operations
A computer stores only bits; any text representation is therefore a mapping from abstract characters onto sequences of numbers called code points. Early seven-bit ASCII proved insufficient for the world’s writing systems, prompting a proliferation of national encodings and, eventually, the Unicode standard. Unicode presently defines more than a million code points and is transmitted most commonly as the variable-length UTF-8 encoding. Python itself stores strings in one of three internal widths—1, 2, or 4 bytes per code point—chosen according to the highest code point present in the string.
Because a single visible character (a grapheme) may be composed of several code points, the built-in len function counts code points rather than user-perceived characters. The rainbow-flag emoji, for example, comprises a white flag, a variation selector, a zero-width joiner, and a rainbow emoji—four code points that render as one glyph. The same phenomenon appears in ordinary text: the Norwegian letter “å” may be stored either as the precomposed code point U+00E5 or as the sequence “a” plus a combining ring (U+030A). Equality tests and slicing that operate on the raw sequence therefore diverge from human expectations.
Case conversion is similarly subtle. The German sharp S (“ß”) upper-cases to “SS”, so a naïve round-trip through str.upper and str.lower fails to restore the original spelling. The Unicode-recommended solution is case folding (str.casefold), which maps characters to a canonical caseless form and correctly handles the 297 known special cases. Even case folding, however, is incomplete for languages such as Turkish, whose dotted and dotless “I” require locale-aware rules that Python does not implement.
Normalization, Security, and Practical Recommendations
Unicode defines four normalization forms. NFC and NFD perform canonical composition and decomposition; NFKC and NFKD additionally map compatibility characters (superscripts, stylistic variants, fractions) onto their plain counterparts. Normalization is essential before comparison or hashing: two strings that look identical to a user may otherwise compare unequal. Python’s identifier parser already applies NFKC, which is why “fancy” mathematical letters are silently rewritten to ordinary ASCII identifiers—an amusing demonstration of the same machinery.
Homoglyphs (characters that look alike but occupy distinct code points) introduce security considerations. The Unicode Consortium publishes confusable lists that applications may consult when validating user names or domain names. Because there is no fixed upper bound on the number of code points that may form a single grapheme cluster, length limits expressed solely in graphemes remain vulnerable to pathological input; a practical defense is to impose both a grapheme limit and a modest code-point ceiling.
For everyday work the speakers recommend a short checklist: always exchange text as UTF-8; compare caselessly with casefold; normalize to a chosen form (usually NFC) before equality tests or storage; treat len and slicing as code-point operations and, when visual length matters, employ a library such as regex or PyICU that understands extended grapheme clusters; and remain aware that Unicode is still evolving and that Python’s support, while extensive, is not exhaustive. Written language is inherently complex; the apparent oddities of Python strings are simply the language’s honest reflection of that complexity.
Links:
[PyConUS2025] Python Software Foundation Welcome and Ecosystem Updates
Lecturer
Deb Nicholson is the Executive Director of the Python Software Foundation. She joined the organization in April 2022 after prior roles at the Open Source Initiative (Interim General Manager), Software Freedom Conservancy (Director of Community Operations), and the Open Invention Network. A founding organizer of the Seattle GNU/Linux Conference, she has received the O’Reilly Open Source Award and the Award for the Advancement of Free Software. Nicholson resides in Cambridge, Massachusetts. Her professional profile is available on LinkedIn at https://www.linkedin.com/in/denicholson and on X under the handle @baconandcoconut.
Abstract
This multi-part welcome session presents the current state of the Python Software Foundation, its infrastructure challenges, security initiatives, and community programs, interwoven with short addresses from major sponsors. Deb Nicholson surveys growth metrics for PyPI, the grants program, staffing, and fiscal sponsorships while reflecting on the social strengths and external pressures facing the community. Accompanying remarks from representatives of AWS, Alpha Omega, Meta, and Google highlight collaborative investments in security, tooling, and language evolution.
Security Initiatives and Infrastructure Stewardship
The session opened with brief sponsor remarks. An NVIDIA representative invited attendees to discussions on faster Python and smaller downloads. Hannah Aubrey of AWS Open Source then framed the company’s commitment to long-term security of critical open-source projects, noting that AWS depends on Python’s reliability. Michael Windsor of Alpha Omega described the organization’s four-pronged approach—staffing dedicated security roles, improving package managers, conducting audits, and funding experiments—initiated in response to supply-chain crises such as Log4Shell. He credited Seth Larson and Mike Fiedler for leadership within the Python ecosystem.
Seth Larson, Security Developer in Residence at the PSF, outlined the contemporary threat landscape: vulnerability exploitation, package poisoning, and social-engineering attacks on maintainers. His mandate is to establish secure-by-default experiences for millions of users. He displayed a visualization of social relationships among top PyPI packages, underscoring that security work is distributed across a dense contributor network rather than concentrated in a single office. Larson invited participation in security-focused talks, open spaces, and a meet-the-experts session at the AWS booth later that day.
Ian Barber of Meta, a visionary sponsor, reported that Python is now the most-used language inside the company, with more than three thousand developers. Meta’s investments include Pyfly, an IDE plugin and type-check engine released in alpha for instant autocomplete and navigation on large codebases; free-threaded Python aimed at true multi-threading without the GIL, particularly beneficial for AI workloads; and continued development of Triton and PyTorch, the latter now supporting NVIDIA’s Blackwell architecture and flex attention. Barber directed attendees to the typing summit and the Meta booth.
Deb Nicholson then assumed the podium. She restated the PSF mission: to promote, protect, and advance the Python language and to foster a diverse international community. The Foundation serves as the nonprofit home for both the language and its community, providing infrastructure, security, and legal support while channeling grants and fiscal sponsorships. Python’s ascent to the most-popular language on GitHub was noted with measured pride, accompanied by gratitude to Fastly for a five-year commitment that underpins CDN capacity.
PyPI growth figures illustrated the scale of demand: an 84 percent rise in download counts and 48 percent rise in bandwidth in 2024, following 57 percent and 45 percent increases in the two preceding years. The newly launched PyPI Organizations feature, intended to support collaborative teams and to generate modest administrative revenue for further improvements, attracted 8 500 initial requests—many duplicates or low-quality. Maria Asha cleared the resulting backlog, and the feature is now fully operational under the stewardship of infrastructure lead E.
The grants program has undergone substantial redesign under Community Communications Manager Marie Nordon. Application processes are shorter, more transparent, and more equitable; community updates are regular. Demand has expanded in every dimension—events, amounts, and geographic reach—while revenue has not kept pace, prompting an open invitation for new funding ideas and corporate partnerships.
Community Resilience, Staffing, and Forward Outlook
Nicholson turned to the social character of the Python community, describing its historic openness to “weird kids” and unconventional projects—ranging from analysis of pre-Columbian tooth enamel to experimental tooling. This receptivity, she argued, keeps the language fresh and expands its application domains. Yet she acknowledged the broader context of geopolitical and economic uncertainty that has deterred hundreds of potential international attendees and complicated nonprofit operations. Layoffs, regulatory flux, and funding volatility place additional strain on a small organization whose largest single budget item remains PyCon US itself.
Despite these pressures, the human infrastructure of the PSF remains robust. Nicholson introduced the approximately thirteen staff members who support more than eight million users worldwide. Olivia Sauls directs the conference; Seth Larson and Mike Fiedler manage security reporting and infrastructure; E oversees core systems with assistance from Jacob Coffee; Marie Nordon handles grants, D&I, fellows, and communications; Jamie has transitioned from contractor to full-time events staff; accountants Phyllis and Laura manage fiscal sponsorships, payroll, and compliance; and Lauren has been promoted to Deputy Executive Director with expanded responsibilities for partnerships and strategic planning. Three CPython Developer-in-Residence positions—Lucas, Peter, and Siri—support core development, with Wukush also present for the language summit.
The Board of Directors and a suite of volunteer working groups (code of conduct, diversity and inclusion, education and outreach, fellows, grants, infrastructure, jobs board, packaging, trademarks) extend capacity further. The Foundation acts as fiscal sponsor for twenty projects, most prominently the global PyLadies network. Thousands of additional volunteers worldwide organize local conferences, meetups, and educational programs that sustain Python’s growth.
Sponsor advocacy received special thanks: the individuals inside corporations who repeatedly champion larger contributions and greater conference attendance. Nicholson closed by urging the community to “stay weird and keep welcoming all the nice weirdos,” then introduced Lisa of Google. Lisa, attending her first PyCon, praised the energy and the dedicated newcomers’ session. She reaffirmed Google’s reliance on Python across scripting, DevOps, and machine learning, and expressed particular interest in free-threaded Python, typing-standard evolution, and performance work. Multiple Google colleagues were present for the typing summit, a talk on programming for oneself, and a supply-chain security open space.
Collectively the addresses portray a Foundation that has scaled its technical and social infrastructure in step with extraordinary adoption while remaining attentive to inclusion, security, and long-term sustainability. The interplay of staff, volunteers, sponsors, and core developers illustrated in the session constitutes the operational backbone that enables the language’s continued vitality.
Links:
[PyConUS2025] Welcome Address by the Conference Chair
Lecturer
Elaine Wong serves as the Conference Chair for PyCon US 2025. A long-standing contributor to the Python community, she previously chaired PyCon Canada in 2018 and co-chaired it in 2019. Wong has organized PyLadies Toronto and the CSV Conference, supported video production at multiple community events, and co-hosts a monthly virtual meetup for conference organizers. Recognized with a Python Software Foundation Community Service Award in 2020 and designated a PSF Fellow, she brings extensive experience in community building, livestreaming, and event logistics. Her professional background includes media and broadcasting work. Further details appear on her personal website at https://eswong.ca/ and her presence on X under the handle @elthenerd.
Abstract
This address opens PyCon US 2025 in Pittsburgh, Pennsylvania, by outlining the logistical framework, community resources, and programmatic highlights of the multi-day gathering. Elaine Wong introduces the organizing team, sponsors, and volunteers while providing essential orientation on venue navigation, code of conduct, accessibility measures, and special tracks. The presentation emphasizes participatory opportunities such as open spaces, hatcheries, and lightning talks, situating the conference within the broader ethos of inclusive Python community practice.
Conference Orientation and Community Infrastructure
Elaine Wong began by expressing enthusiasm for the assembled attendees and offering a light-hearted musical greeting adapted to the occasion. She introduced herself as Conference Chair and acknowledged her co-chair, John, who stood ready to assist with questions throughout the event. Together they form the primary points of contact for logistical support and guidance.
Central to the success of the conference, Wong stressed, is the dedicated staff of the Python Software Foundation. She listed Olivia, Deb, Lauren, Jamie, Phyllis, Laura, E, Marie, Jacob, Maria, Seth, and Mike, inviting a round of applause for their year-long efforts in preparing the venue, program, and infrastructure. Equally vital are the sponsors whose financial contributions underwrite food service, utilities, and overall operations; attendees were encouraged to visit the expo hall and express personal thanks via a provided QR code linking to the full sponsor roster.
The volunteer cohort received particular recognition. More than one hundred individuals had already signed up, with a target of three hundred on-site volunteer hours. Wong directed participants to a QR code and website for short-term shifts ranging from one to several hours, noting that such engagement offers both practical insight into conference operations and an eventual pathway toward leadership roles such as chair.
Practical orientation followed. Wi-Fi credentials and a QR code for network access were displayed, alongside announcement of a new mobile application available for both iOS and Android platforms. The application supplies schedules, maps, and session favoriting capabilities. Venue layout received detailed treatment: general sessions, keynotes, lightning talks, and special guest presentations occupy Hall B on the second floor. Parallel talk tracks are located in Hall C and on the third floor, reachable via escalators or elevators and a pedestrian walkway. Breakfast is served outside Hall B beginning at 7:30 each morning; lunch and snacks appear in Hall A, which also houses the expo hall on Friday and Saturday as well as swag pickup for pre-ordered t-shirts. Surplus shirts become available for purchase on Sunday at 11:00 on a first-come basis. Luggage storage operates Saturday and Sunday outside Show Office C from 8:00 to 18:00, with later retrieval possible from the staff office in rooms 306–307. A mothering and nursing room occupies Show Office B; the travel-grant office is in Show Office A, with posted hours for recipients. Gender-neutral restrooms are clearly marked. The fourth floor hosts summits and the PyLadies lunch, while room 405 provides a quiet space and a rooftop terrace offers outdoor respite. Emergency procedures direct participants to call 911 and notify staff wearing light-blue shirts or DLCC attire; on-site EMTs are stationed behind Hall B on the main conference days.
Code-of-conduct provisions received careful attention. Wong summarized the expected standard of conduct by reference to the kindness and consideration associated with Mr. Rogers. Concerns may be reported by text, email, or in person to the committee led by Molly, whose members wear distinctive orange shirts and maintain a presence in the staff room. Photo policy is signaled by lanyard color: orange indicates preference against photography, black indicates consent. Health and safety guidance encourages but does not require masking; free masks and testing kits remain available at registration.
A brief recap of the preceding two days noted completed tutorials, the language summit, education summit, WebAssembly summit, sponsored presentations, newcomers’ orientation, and a lively opening reception. Keynote speakers Jeff, Tom, Lynn, Corey, and Dr. Corey Jordan were introduced, with the added note that all would hold meet-and-greet sessions at the PSF booth; the Carpentries, led by Dr. Jordan, maintain an expo presence. Special programming includes a diversity-and-inclusion panel, a security-engineer update, and a Python Steering Council session. The Charlas Spanish-language track occupies rooms 310 and 311 for two days, inviting participants of every proficiency level. Live human captioning, provided by White Coat Captioning, accompanies all talks.
Hatchery events expand the formal program: FlaskCon occupies the afternoon of the current day, a community-organizer summit and Hometown Heroes showcase occur on Saturday, and a beginner-oriented Humble Data workshop runs on Sunday, all centered in room 317. Additional summits address typing, maintainership, packaging, and mentor sprints. Open spaces have been digitized; participants register via website or QR code for one-hour self-organized sessions on a first-come basis. Lightning-talk sign-ups likewise proceed through an online form linked to each attendee’s dashboard, with four scheduled rounds. Tickets remain available for the PyLadies auction, a dinner-and-bidding event that constitutes the principal fundraiser for the PyLadies Foundation and supports both global chapters and travel grants. The PyLadies luncheon on Sunday offers further reflection and networking. Sunday also features posters, a job fair, and a community showcase that invites regional conferences, meetups, and open-source projects to claim tables.
Local recommendations for Pittsburgh attractions, including the baseball stadium and neighborhood eateries, appear under a provided URL. Wong closed by endorsing the informal “hallway track” and the Pac-Man rule—leaving physical space in conversational circles so that newcomers may join readily. Official social-media channels and the hashtag #PyConUS complete the orientation.
Programmatic Opportunities and Inclusive Practices
Beyond logistics, the address situates PyCon US within a deliberately inclusive framework. The multi-track structure, language accessibility via Charlas and live captioning, quiet rooms, gender-neutral facilities, and explicit photo and health policies collectively lower barriers to participation. Hatcheries and open spaces democratize content creation, allowing attendees to surface emerging ideas outside the formal call-for-proposals process. Lightning talks lower the threshold for first-time speakers. Volunteer pathways and the explicit invitation to future leadership roles reinforce the conference’s self-renewing character. Financial underwriting by sponsors and the travel-grant system further extend reach. Collectively these measures illustrate how a large-scale technical gathering can embed community values of accessibility, mutual support, and continuous expansion of the Python ecosystem.
Links:
Demystifying Parquet: The Power of Efficient Data Storage in the Cloud
Unlocking the Power of Apache Parquet: A Modern Standard for Data Efficiency
In today’s digital ecosystem, where data volume, velocity, and variety continue to rise, the choice of file format can dramatically impact performance, scalability, and cost. Whether you are an architect designing a cloud-native data platform or a developer managing analytics pipelines, Apache Parquet stands out as a foundational technology you should understand — and probably already rely on.
This article explores what Parquet is, why it matters, and how to work with it in practice — including real examples in Python, Java, Node.js, and Bash for converting and uploading files to Amazon S3.
What Is Apache Parquet?
Apache Parquet is a high-performance, open-source file format designed for efficient columnar data storage. Originally developed by Twitter and Cloudera and now an Apache Software Foundation project, Parquet is purpose-built for use with distributed data processing frameworks like Apache Spark, Hive, Impala, and Drill.
Unlike row-based formats such as CSV or JSON, Parquet organizes data by columns rather than rows. This enables powerful compression, faster retrieval of selected fields, and dramatic performance improvements for analytical queries.
Why Choose Parquet?
✅ Columnar Format = Faster Queries
Because Parquet stores values from the same column together, analytical engines can skip irrelevant data and process only what’s required — reducing I/O and boosting speed.
Compression and Storage Efficiency
Parquet achieves better compression ratios than row-based formats, thanks to the similarity of values in each column. This translates directly into reduced cloud storage costs.
Schema Evolution
Parquet supports schema evolution, enabling your datasets to grow gracefully. New fields can be added over time without breaking existing consumers.
Interoperability
The format is compatible across multiple ecosystems and languages, including Python (Pandas, PyArrow), Java (Spark, Hadoop), and even browser-based analytics tools.
☁️ Using Parquet with Amazon S3
One of the most common modern use cases for Parquet is in conjunction with Amazon S3, where it powers data lakes, ETL pipelines, and serverless analytics via services like Amazon Athena and Redshift Spectrum.
Here’s how you can write Parquet files and upload them to S3 in different environments:
From CSV to Parquet in Practice
Python Example
import pandas as pd
# Load CSV data
df = pd.read_csv("input.csv")
# Save as Parquet
df.to_parquet("output.parquet", engine="pyarrow")
To upload to S3:
import boto3
s3 = boto3.client("s3")
s3.upload_file("output.parquet", "your-bucket", "data/output.parquet")
Node.js Example
Install the required libraries:
npm install aws-sdk
Upload file to S3:
const AWS = require('aws-sdk');
const fs = require('fs');
const s3 = new AWS.S3();
const fileContent = fs.readFileSync('output.parquet');
const params = {
Bucket: 'your-bucket',
Key: 'data/output.parquet',
Body: fileContent
};
s3.upload(params, (err, data) => {
if (err) throw err;
console.log(`File uploaded successfully at ${data.Location}`);
});
☕ Java with Apache Spark and AWS SDK
In your pom.xml, include:
<dependency>
<groupId>org.apache.parquet</groupId>
<artifactId>parquet-hadoop</artifactId>
<version>1.12.2</version>
</dependency>
<dependency>
<groupId>com.amazonaws</groupId>
<artifactId>aws-java-sdk-s3</artifactId>
<version>1.12.470</version>
</dependency>
Spark conversion:
Dataset<Row> df = spark.read().option("header", "true").csv("input.csv");
df.write().parquet("output.parquet");
Upload to S3:
AmazonS3 s3 = AmazonS3ClientBuilder.standard()
.withRegion("us-west-2")
.withCredentials(new AWSStaticCredentialsProvider(
new BasicAWSCredentials("ACCESS_KEY", "SECRET_KEY")))
.build();
s3.putObject("your-bucket", "data/output.parquet", new File("output.parquet"));
Bash with AWS CLI
aws s3 cp output.parquet s3://your-bucket/data/output.parquet
Final Thoughts
Apache Parquet has quietly become a cornerstone of the modern data stack. It powers everything from ad hoc analytics to petabyte-scale data lakes, bringing consistency and efficiency to how we store and retrieve data.
Whether you are migrating legacy pipelines, designing new AI workloads, or simply optimizing your storage bills — understanding and adopting Parquet can unlock meaningful benefits.
When used in combination with cloud platforms like AWS, the performance, scalability, and cost-efficiency of Parquet-based workflows are hard to beat.
Creating EPUBs from Images: A Developer’s Guide to Digital Publishing
Ever needed to convert a collection of images into a professional EPUB file? Whether you’re working with comics, manga, or any image-based content, I’ve developed a Python script that makes this process seamless and customizable.
What is create_epub.py?
This Python script transforms a folder of images into a fully-featured EPUB file, complete with:
- Proper EPUB 3.0 structure
- Customizable metadata
- Table of contents
- Responsive image display
- Cover image handling
Key Features
- Smart Filename Generation: Automatically generates EPUB filenames based on metadata (e.g., “MyBook_01_1.epub”)
- Comprehensive Metadata Support: Title, author, series, volume, edition, ISBN, and more
- Image Optimization: Supports JPEG, PNG, and GIF formats with proper scaling
- Responsive Design: CSS-based layout that works across devices
- Detailed Logging: Progress tracking and debugging capabilities
Usage Example
python create_epub.py image_folder \
--title "My Book" \
--author "Author Name" \
--volume 1 \
--edition "First Edition" \
--series "My Series" \
--publisher "My Publisher" \
--isbn "978-3-16-148410-0"
Technical Details
The script creates a proper EPUB 3.0 structure with:
- META-INF/container.xml
- OEBPS/content.opf (metadata)
- OEBPS/toc.ncx (table of contents)
- OEBPS/nav.xhtml (navigation)
- OEBPS/style.css (responsive styling)
- OEBPS/images/ (image storage)
Best Practices Implemented
- Proper XML namespaces and validation
- Responsive image handling
- Comprehensive metadata support
- Clean, maintainable code structure
- Extensive error handling and logging
Getting Started
# Install dependencies
pip install -r requirements.txt
# Basic usage
python create_epub.py /path/to/images --title "My Book"
# With debug logging
python create_epub.py /path/to/images --title "My Book" --debug
The script is designed to be both powerful and user-friendly, making it accessible to developers while providing the flexibility needed for professional publishing workflows.
Whether you’re a developer looking to automate EPUB creation or a content creator seeking to streamline your publishing process, this tool provides a robust solution for converting images into EPUB files.
The script on GitHub or below: 👇👇👇
[python]
import os
import sys
import logging
import zipfile
import uuid
from datetime import datetime
import argparse
from PIL import Image
import xml.etree.ElementTree
from xml.dom import minidom
# @author Jonathan Lalou / https://github.com/JonathanLalou/
# Configure logging
logging.basicConfig(
level=logging.INFO,
format=’%(asctime)s – %(levelname)s – %(message)s’,
handlers=[
logging.StreamHandler(sys.stdout)
]
)
logger = logging.getLogger(__name__)
# Define the CSS content
CSS_CONTENT = ”’
body {
margin: 0;
padding: 0;
display: flex;
justify-content: center;
align-items: center;
min-height: 100vh;
}
img {
max-width: 100%;
max-height: 100vh;
object-fit: contain;
}
”’
def create_container_xml():
"""Create the container.xml file."""
logger.debug("Creating container.xml")
container = xml.etree.ElementTree.Element(‘container’, {
‘version’: ‘1.0’,
‘xmlns’: ‘urn:oasis:names:tc:opendocument:xmlns:container’
})
rootfiles = xml.etree.ElementTree.SubElement(container, ‘rootfiles’)
xml.etree.ElementTree.SubElement(rootfiles, ‘rootfile’, {
‘full-path’: ‘OEBPS/content.opf’,
‘media-type’: ‘application/oebps-package+xml’
})
xml_content = prettify_xml(container)
logger.debug("container.xml content:\n" + xml_content)
return xml_content
def create_content_opf(metadata, spine_items, manifest_items):
"""Create the content.opf file."""
logger.debug("Creating content.opf")
logger.debug(f"Metadata: {metadata}")
logger.debug(f"Spine items: {spine_items}")
logger.debug(f"Manifest items: {manifest_items}")
package = xml.etree.ElementTree.Element(‘package’, {
‘xmlns’: ‘http://www.idpf.org/2007/opf’,
‘xmlns:dc’: ‘http://purl.org/dc/elements/1.1/’,
‘xmlns:dcterms’: ‘http://purl.org/dc/terms/’,
‘xmlns:opf’: ‘http://www.idpf.org/2007/opf’,
‘version’: ‘3.0’,
‘unique-identifier’: ‘bookid’
})
# Metadata
metadata_elem = xml.etree.ElementTree.SubElement(package, ‘metadata’)
# Required metadata
book_id = str(uuid.uuid4())
xml.etree.ElementTree.SubElement(metadata_elem, ‘dc:identifier’, {‘id’: ‘bookid’}).text = book_id
logger.debug(f"Generated book ID: {book_id}")
xml.etree.ElementTree.SubElement(metadata_elem, ‘dc:title’).text = metadata.get(‘title’, ‘Untitled’)
xml.etree.ElementTree.SubElement(metadata_elem, ‘dc:language’).text = metadata.get(‘language’, ‘en’)
xml.etree.ElementTree.SubElement(metadata_elem, ‘dc:creator’).text = metadata.get(‘author’, ‘Unknown’)
# Add required dcterms:modified
current_time = datetime.now().strftime(‘%Y-%m-%dT%H:%M:%SZ’)
xml.etree.ElementTree.SubElement(metadata_elem, ‘meta’, {
‘property’: ‘dcterms:modified’
}).text = current_time
# Add cover metadata
xml.etree.ElementTree.SubElement(metadata_elem, ‘meta’, {
‘name’: ‘cover’,
‘content’: ‘cover-image’
})
# Add additional metadata
if metadata.get(‘publisher’):
xml.etree.ElementTree.SubElement(metadata_elem, ‘dc:publisher’).text = metadata[‘publisher’]
if metadata.get(‘description’):
xml.etree.ElementTree.SubElement(metadata_elem, ‘dc:description’).text = metadata[‘description’]
if metadata.get(‘rights’):
xml.etree.ElementTree.SubElement(metadata_elem, ‘dc:rights’).text = metadata[‘rights’]
if metadata.get(‘subject’):
xml.etree.ElementTree.SubElement(metadata_elem, ‘dc:subject’).text = metadata[‘subject’]
if metadata.get(‘isbn’):
xml.etree.ElementTree.SubElement(metadata_elem, ‘dc:identifier’, {
‘opf:scheme’: ‘ISBN’
}).text = metadata[‘isbn’]
# Series metadata
if metadata.get(‘series’):
xml.etree.ElementTree.SubElement(metadata_elem, ‘meta’, {
‘property’: ‘belongs-to-collection’
}).text = metadata[‘series’]
xml.etree.ElementTree.SubElement(metadata_elem, ‘meta’, {
‘property’: ‘group-position’
}).text = metadata.get(‘volume’, ‘1’)
# Release date
if metadata.get(‘release_date’):
xml.etree.ElementTree.SubElement(metadata_elem, ‘dc:date’).text = metadata[‘release_date’]
# Version and edition
if metadata.get(‘version’):
xml.etree.ElementTree.SubElement(metadata_elem, ‘meta’, {
‘property’: ‘schema:version’
}).text = metadata[‘version’]
if metadata.get(‘edition’):
xml.etree.ElementTree.SubElement(metadata_elem, ‘meta’, {
‘property’: ‘schema:bookEdition’
}).text = metadata[‘edition’]
# Manifest
manifest = xml.etree.ElementTree.SubElement(package, ‘manifest’)
for item in manifest_items:
xml.etree.ElementTree.SubElement(manifest, ‘item’, item)
# Spine
spine = xml.etree.ElementTree.SubElement(package, ‘spine’)
for item in spine_items:
xml.etree.ElementTree.SubElement(spine, ‘itemref’, {‘idref’: item})
xml_content = prettify_xml(package)
logger.debug("content.opf content:\n" + xml_content)
return xml_content
def create_toc_ncx(metadata, nav_points):
"""Create the toc.ncx file."""
logger.debug("Creating toc.ncx")
logger.debug(f"Navigation points: {nav_points}")
ncx = xml.etree.ElementTree.Element(‘ncx’, {
‘xmlns’: ‘http://www.daisy.org/z3986/2005/ncx/’,
‘version’: ‘2005-1’
})
head = xml.etree.ElementTree.SubElement(ncx, ‘head’)
book_id = str(uuid.uuid4())
xml.etree.ElementTree.SubElement(head, ‘meta’, {‘name’: ‘dtb:uid’, ‘content’: book_id})
logger.debug(f"Generated NCX book ID: {book_id}")
xml.etree.ElementTree.SubElement(head, ‘meta’, {‘name’: ‘dtb:depth’, ‘content’: ‘1’})
xml.etree.ElementTree.SubElement(head, ‘meta’, {‘name’: ‘dtb:totalPageCount’, ‘content’: ‘0’})
xml.etree.ElementTree.SubElement(head, ‘meta’, {‘name’: ‘dtb:maxPageNumber’, ‘content’: ‘0’})
doc_title = xml.etree.ElementTree.SubElement(ncx, ‘docTitle’)
xml.etree.ElementTree.SubElement(doc_title, ‘text’).text = metadata.get(‘title’, ‘Untitled’)
nav_map = xml.etree.ElementTree.SubElement(ncx, ‘navMap’)
for i, (id, label, src) in enumerate(nav_points, 1):
nav_point = xml.etree.ElementTree.SubElement(nav_map, ‘navPoint’, {‘id’: id, ‘playOrder’: str(i)})
nav_label = xml.etree.ElementTree.SubElement(nav_point, ‘navLabel’)
xml.etree.ElementTree.SubElement(nav_label, ‘text’).text = label
xml.etree.ElementTree.SubElement(nav_point, ‘content’, {‘src’: src})
xml_content = prettify_xml(ncx)
logger.debug("toc.ncx content:\n" + xml_content)
return xml_content
def create_nav_xhtml(metadata, nav_points):
"""Create the nav.xhtml file."""
logger.debug("Creating nav.xhtml")
html = xml.etree.ElementTree.Element(‘html’, {
‘xmlns’: ‘http://www.w3.org/1999/xhtml’,
‘xmlns:epub’: ‘http://www.idpf.org/2007/ops’
})
head = xml.etree.ElementTree.SubElement(html, ‘head’)
xml.etree.ElementTree.SubElement(head, ‘title’).text = ‘Table of Contents’
body = xml.etree.ElementTree.SubElement(html, ‘body’)
nav = xml.etree.ElementTree.SubElement(body, ‘nav’, {‘epub:type’: ‘toc’})
ol = xml.etree.ElementTree.SubElement(nav, ‘ol’)
for _, label, src in nav_points:
li = xml.etree.ElementTree.SubElement(ol, ‘li’)
xml.etree.ElementTree.SubElement(li, ‘a’, {‘href’: src}).text = label
xml_content = prettify_xml(html)
logger.debug("nav.xhtml content:\n" + xml_content)
return xml_content
def create_page_xhtml(page_number, image_file):
"""Create an XHTML page for an image."""
logger.debug(f"Creating page {page_number} for image {image_file}")
html = xml.etree.ElementTree.Element(‘html’, {
‘xmlns’: ‘http://www.w3.org/1999/xhtml’,
‘xmlns:epub’: ‘http://www.idpf.org/2007/ops’
})
head = xml.etree.ElementTree.SubElement(html, ‘head’)
xml.etree.ElementTree.SubElement(head, ‘title’).text = f’Page {page_number}’
xml.etree.ElementTree.SubElement(head, ‘link’, {
‘rel’: ‘stylesheet’,
‘type’: ‘text/css’,
‘href’: ‘style.css’
})
body = xml.etree.ElementTree.SubElement(html, ‘body’)
xml.etree.ElementTree.SubElement(body, ‘img’, {
‘src’: f’images/{image_file}’,
‘alt’: f’Page {page_number}’
})
xml_content = prettify_xml(html)
logger.debug(f"Page {page_number} XHTML content:\n" + xml_content)
return xml_content
def prettify_xml(elem):
"""Convert XML element to pretty string."""
rough_string = xml.etree.ElementTree.tostring(elem, ‘utf-8’)
reparsed = minidom.parseString(rough_string)
return reparsed.toprettyxml(indent=" ")
def create_epub_from_images(image_folder, output_file, metadata):
logger.info(f"Starting EPUB creation from images in {image_folder}")
logger.info(f"Output file will be: {output_file}")
logger.info(f"Metadata: {metadata}")
# Get all image files
image_files = [f for f in os.listdir(image_folder)
if f.lower().endswith((‘.png’, ‘.jpg’, ‘.jpeg’, ‘.gif’, ‘.bmp’))]
image_files.sort()
logger.info(f"Found {len(image_files)} image files")
logger.debug(f"Image files: {image_files}")
if not image_files:
logger.error("No image files found in the specified folder")
sys.exit(1)
# Create ZIP file (EPUB)
logger.info("Creating EPUB file structure")
with zipfile.ZipFile(output_file, ‘w’, zipfile.ZIP_DEFLATED) as epub:
# Add mimetype (must be first, uncompressed)
logger.debug("Adding mimetype file (uncompressed)")
epub.writestr(‘mimetype’, ‘application/epub+zip’, zipfile.ZIP_STORED)
# Create META-INF directory
logger.debug("Adding container.xml")
epub.writestr(‘META-INF/container.xml’, create_container_xml())
# Create OEBPS directory structure
logger.debug("Creating OEBPS directory structure")
os.makedirs(‘temp/OEBPS/images’, exist_ok=True)
os.makedirs(‘temp/OEBPS/style’, exist_ok=True)
# Add CSS
logger.debug("Adding style.css")
epub.writestr(‘OEBPS/style.css’, CSS_CONTENT)
# Process images and create pages
logger.info("Processing images and creating pages")
manifest_items = [
{‘id’: ‘style’, ‘href’: ‘style.css’, ‘media-type’: ‘text/css’},
{‘id’: ‘nav’, ‘href’: ‘nav.xhtml’, ‘media-type’: ‘application/xhtml+xml’, ‘properties’: ‘nav’}
]
spine_items = []
nav_points = []
for i, image_file in enumerate(image_files, 1):
logger.debug(f"Processing image {i:03d}/{len(image_files):03d}: {image_file}")
# Copy image to temp directory
image_path = os.path.join(image_folder, image_file)
logger.debug(f"Reading image: {image_path}")
with open(image_path, ‘rb’) as f:
image_data = f.read()
logger.debug(f"Adding image to EPUB: OEBPS/images/{image_file}")
epub.writestr(f’OEBPS/images/{image_file}’, image_data)
# Add image to manifest
image_id = f’image_{i:03d}’
if i == 1:
image_id = ‘cover-image’ # Special ID for cover image
manifest_items.append({
‘id’: image_id,
‘href’: f’images/{image_file}’,
‘media-type’: ‘image/jpeg’ if image_file.lower().endswith((‘.jpg’, ‘.jpeg’)) else ‘image/png’
})
# Create page XHTML
page_id = f’page_{i:03d}’
logger.debug(f"Creating page XHTML: {page_id}.xhtml")
page_content = create_page_xhtml(i, image_file)
epub.writestr(f’OEBPS/{page_id}.xhtml’, page_content)
# Add to manifest and spine
manifest_items.append({
‘id’: page_id,
‘href’: f'{page_id}.xhtml’,
‘media-type’: ‘application/xhtml+xml’
})
spine_items.append(page_id)
# Add to navigation points
nav_points.append((
f’navpoint-{i:03d}’,
‘Cover’ if i == 1 else f’Page {i:03d}’,
f'{page_id}.xhtml’
))
# Create content.opf
logger.debug("Creating content.opf")
epub.writestr(‘OEBPS/content.opf’, create_content_opf(metadata, spine_items, manifest_items))
# Create toc.ncx
logger.debug("Creating toc.ncx")
epub.writestr(‘OEBPS/toc.ncx’, create_toc_ncx(metadata, nav_points))
# Create nav.xhtml
logger.debug("Creating nav.xhtml")
epub.writestr(‘OEBPS/nav.xhtml’, create_nav_xhtml(metadata, nav_points))
logger.info(f"Successfully created EPUB file: {output_file}")
logger.info("EPUB structure:")
logger.info(" mimetype")
logger.info(" META-INF/container.xml")
logger.info(" OEBPS/")
logger.info(" content.opf")
logger.info(" toc.ncx")
logger.info(" nav.xhtml")
logger.info(" style.css")
logger.info(" images/")
for i in range(1, len(image_files) + 1):
logger.info(f" page_{i:03d}.xhtml")
def generate_default_filename(metadata, image_folder):
"""Generate default EPUB filename based on metadata."""
# Get title from metadata or use folder name
title = metadata.get(‘title’)
if not title:
# Get folder name and extract part before last underscore
folder_name = os.path.basename(os.path.normpath(image_folder))
title = folder_name.rsplit(‘_’, 1)[0] if ‘_’ in folder_name else folder_name
# Format title: remove spaces, hyphens, quotes and capitalize
title = ”.join(word.capitalize() for word in title.replace(‘-‘, ‘ ‘).replace(‘"’, ”).replace("’", ”).split())
# Format volume number with 2 digits
volume = metadata.get(‘volume’, ’01’)
if volume.isdigit():
volume = f"{int(volume):02d}"
# Get edition number
edition = metadata.get(‘edition’, ‘1’)
return f"{title}_{volume}_{edition}.epub"
def main():
parser = argparse.ArgumentParser(description=’Create an EPUB from a folder of images’)
parser.add_argument(‘image_folder’, help=’Folder containing the images’)
parser.add_argument(‘–output-file’, ‘-o’, help=’Output EPUB file path (optional)’)
parser.add_argument(‘–title’, help=’Book title’)
parser.add_argument(‘–author’, help=’Book author’)
parser.add_argument(‘–series’, help=’Series name’)
parser.add_argument(‘–volume’, help=’Volume number’)
parser.add_argument(‘–release-date’, help=’Release date (YYYY-MM-DD)’)
parser.add_argument(‘–edition’, help=’Edition number’)
parser.add_argument(‘–version’, help=’Version number’)
parser.add_argument(‘–language’, help=’Book language (default: en)’)
parser.add_argument(‘–publisher’, help=’Publisher name’)
parser.add_argument(‘–description’, help=’Book description’)
parser.add_argument(‘–rights’, help=’Copyright/license information’)
parser.add_argument(‘–subject’, help=’Book subject/category’)
parser.add_argument(‘–isbn’, help=’ISBN number’)
parser.add_argument(‘–debug’, action=’store_true’, help=’Enable debug logging’)
args = parser.parse_args()
if args.debug:
logger.setLevel(logging.DEBUG)
logger.info("Debug logging enabled")
if not os.path.exists(args.image_folder):
logger.error(f"Image folder does not exist: {args.image_folder}")
sys.exit(1)
if not os.path.isdir(args.image_folder):
logger.error(f"Specified path is not a directory: {args.image_folder}")
sys.exit(1)
metadata = {
‘title’: args.title,
‘author’: args.author,
‘series’: args.series,
‘volume’: args.volume,
‘release_date’: args.release_date,
‘edition’: args.edition,
‘version’: args.version,
‘language’: args.language,
‘publisher’: args.publisher,
‘description’: args.description,
‘rights’: args.rights,
‘subject’: args.subject,
‘isbn’: args.isbn
}
# Remove None values from metadata
metadata = {k: v for k, v in metadata.items() if v is not None}
# Generate output filename if not provided
if not args.output_file:
args.output_file = generate_default_filename(metadata, args.image_folder)
logger.info(f"Using default output filename: {args.output_file}")
try:
create_epub_from_images(args.image_folder, args.output_file, metadata)
logger.info("EPUB creation completed successfully")
except Exception as e:
logger.error(f"EPUB creation failed: {str(e)}")
sys.exit(1)
if __name__ == ‘__main__’:
main()
[/python]
Understanding Chi-Square Tests: A Comprehensive Guide for Developers
In the world of software development and data analysis, understanding statistical significance is crucial. Whether you’re running A/B tests, analyzing user behavior, or building machine learning models, the Chi-Square (χ²) test is an essential tool in your statistical toolkit. This comprehensive guide will help you understand its principles, implementation, and practical applications.
What is Chi-Square?
The Chi-Square test is a statistical method used to determine if there’s a significant difference between expected and observed frequencies in categorical data. It’s named after the Greek letter χ (chi) and is particularly useful for analyzing relationships between categorical variables.
Historical Context
The Chi-Square test was developed by Karl Pearson in 1900, making it one of the oldest statistical tests still in widespread use today. Its development marked a significant advancement in statistical analysis, particularly in the field of categorical data analysis.
Core Principles and Mathematical Foundation
- Null Hypothesis (H₀): Assumes no significant difference between observed and expected data
- Alternative Hypothesis (H₁): Suggests a significant difference exists
- Degrees of Freedom: Number of categories minus constraints
- P-value: Probability of observing the results if H₀ is true
The Chi-Square Formula
The Chi-Square statistic is calculated using the formula:
χ² = Σ [(O - E)² / E]
Where: – O = Observed frequency – E = Expected frequency – Σ = Sum over all categories
Practical Implementation
1. A/B Testing Implementation (Python)
from scipy.stats import chi2_contingency
import numpy as np
import matplotlib.pyplot as plt
def perform_ab_test(control_data, treatment_data):
"""
Perform A/B test using Chi-Square test
Args:
control_data: List of [successes, failures] for control group
treatment_data: List of [successes, failures] for treatment group
"""
# Create contingency table
observed = np.array([control_data, treatment_data])
# Perform Chi-Square test
chi2, p_value, dof, expected = chi2_contingency(observed)
# Calculate effect size (Cramer's V)
n = np.sum(observed)
min_dim = min(observed.shape) - 1
cramers_v = np.sqrt(chi2 / (n * min_dim))
return {
'chi2': chi2,
'p_value': p_value,
'dof': dof,
'expected': expected,
'effect_size': cramers_v
}
# Example usage
control = [100, 150] # [clicks, no-clicks] for control
treatment = [120, 130] # [clicks, no-clicks] for treatment
results = perform_ab_test(control, treatment)
print(f"Chi-Square: {results['chi2']:.2f}")
print(f"P-value: {results['p_value']:.4f}")
print(f"Effect Size (Cramer's V): {results['effect_size']:.3f}")
2. Feature Selection Implementation (Java)
import org.apache.commons.math3.stat.inference.ChiSquareTest;
import java.util.Arrays;
public class FeatureSelection {
private final ChiSquareTest chiSquareTest;
public FeatureSelection() {
this.chiSquareTest = new ChiSquareTest();
}
public FeatureSelectionResult analyzeFeature(
long[][] observed,
double significanceLevel) {
double pValue = chiSquareTest.chiSquareTest(observed);
boolean isSignificant = pValue < significanceLevel;
// Calculate effect size (Cramer's V)
double chiSquare = chiSquareTest.chiSquare(observed);
long total = Arrays.stream(observed)
.flatMapToLong(Arrays::stream)
.sum();
int minDim = Math.min(observed.length, observed[0].length) - 1;
double cramersV = Math.sqrt(chiSquare / (total * minDim));
return new FeatureSelectionResult(
pValue,
isSignificant,
cramersV
);
}
public static class FeatureSelectionResult {
private final double pValue;
private final boolean isSignificant;
private final double effectSize;
// Constructor and getters
}
}
Advanced Applications
1. Machine Learning Feature Selection
Chi-Square tests are particularly useful in feature selection for machine learning models. Here’s how to implement it in Python using scikit-learn:
from sklearn.feature_selection import SelectKBest, chi2
from sklearn.datasets import load_iris
import pandas as pd
# Load dataset
iris = load_iris()
X = pd.DataFrame(iris.data, columns=iris.feature_names)
y = iris.target
# Select top 2 features using Chi-Square
selector = SelectKBest(chi2, k=2)
X_new = selector.fit_transform(X, y)
# Get selected features
selected_features = X.columns[selector.get_support()]
print(f"Selected features: {selected_features.tolist()}")
2. Goodness-of-Fit Testing
Testing if your data follows a particular distribution:
from scipy.stats import chisquare
import numpy as np
# Example: Testing if dice is fair
observed = np.array([18, 16, 15, 17, 16, 18]) # Observed frequencies
expected = np.array([16.67, 16.67, 16.67, 16.67, 16.67, 16.67]) # Expected for fair dice
chi2, p_value = chisquare(observed, expected)
print(f"Chi-Square: {chi2:.2f}")
print(f"P-value: {p_value:.4f}")
Best Practices and Considerations
- Sample Size: Ensure sufficient sample size for reliable results
- Expected Frequencies: Each expected frequency should be ≥ 5
- Multiple Testing: Apply corrections (e.g., Bonferroni) when conducting multiple tests
- Effect Size: Consider effect size in addition to p-values
- Assumptions: Verify test assumptions before application
Common Pitfalls to Avoid
- Using Chi-Square for continuous data
- Ignoring small expected frequencies
- Overlooking multiple testing issues
- Focusing solely on p-values without considering effect size
- Applying the test without checking assumptions
Resources and Further Reading
- Scipy Chi-Square Documentation
- Apache Commons Math
- Interactive Chi-Square Calculator
- Wikipedia: Chi-Squared Test
Understanding and properly implementing Chi-Square tests can significantly enhance your data analysis capabilities as a developer. Whether you’re working on A/B testing, feature selection, or data validation, this statistical tool provides valuable insights into your data’s relationships and distributions.
Remember to always consider the context of your analysis, verify assumptions, and interpret results carefully. Happy coding!
RSS to EPUB Converter: Create eBooks from RSS Feeds
Overview
This Python script (rss_to_ebook.py) converts RSS or Atom feeds into EPUB format eBooks, allowing you to read your favorite blog posts and news articles offline in your preferred e-reader. The script intelligently handles both RSS 2.0 and Atom feed formats, preserving HTML formatting while creating a clean, readable eBook.
Key Features
- Dual Format Support: Works with both RSS 2.0 and Atom feeds
- Smart Pagination: Automatically handles paginated feeds using multiple detection methods
- Date Range Filtering: Select specific date ranges for content inclusion
- Metadata Preservation: Maintains feed metadata including title, author, and description
- HTML Formatting: Preserves original HTML formatting while cleaning unnecessary elements
- Duplicate Prevention: Automatically detects and removes duplicate entries
- Comprehensive Logging: Detailed progress tracking and error reporting
Technical Details
The script uses several Python libraries:
feedparser: For parsing RSS and Atom feedsebooklib: For creating EPUB filesBeautifulSoup: For HTML cleaning and processinglogging: For detailed operation tracking
Usage
python rss_to_ebook.py <feed_url> [--start-date YYYY-MM-DD] [--end-date YYYY-MM-DD] [--output filename.epub] [--debug]
Parameters:
feed_url: URL of the RSS or Atom feed (required)--start-date: Start date for content inclusion (default: 1 year ago)--end-date: End date for content inclusion (default: today)--output: Output EPUB filename (default: rss_feed.epub)--debug: Enable detailed logging
Example
python rss_to_ebook.py https://example.com/feed --start-date 2024-01-01 --end-date 2024-03-31 --output my_blog.epub
Requirements
- Python 3.x
- Required packages (install via pip):
pip install feedparser ebooklib beautifulsoup4
How It Works
- Feed Detection: Automatically identifies feed format (RSS 2.0 or Atom)
- Content Processing:
- Extracts entries within specified date range
- Preserves HTML formatting while cleaning unnecessary elements
- Handles pagination to get all available content
- EPUB Creation:
- Creates chapters from feed entries
- Maintains original formatting and links
- Includes table of contents and navigation
- Preserves feed metadata
Error Handling
- Validates feed format and content
- Handles malformed HTML
- Provides detailed error messages and logging
- Gracefully handles missing or incomplete feed data
Use Cases
- Create eBooks from your favorite blogs
- Archive important news articles
- Generate reading material for offline use
- Create compilations of related content
Gist: GitHub
Here is the script:
[python]
#!/usr/bin/env python3
import feedparser
import argparse
from datetime import datetime, timedelta
from ebooklib import epub
import re
from bs4 import BeautifulSoup
import logging
# Configure logging
logging.basicConfig(
level=logging.INFO,
format=’%(asctime)s – %(levelname)s – %(message)s’,
datefmt=’%Y-%m-%d %H:%M:%S’
)
def clean_html(html_content):
"""Clean HTML content while preserving formatting."""
soup = BeautifulSoup(html_content, ‘html.parser’)
# Remove script and style elements
for script in soup(["script", "style"]):
script.decompose()
# Remove any inline styles
for tag in soup.find_all(True):
if ‘style’ in tag.attrs:
del tag.attrs[‘style’]
# Return the cleaned HTML
return str(soup)
def get_next_feed_page(current_feed, feed_url):
"""Get the next page of the feed using various pagination methods."""
# Method 1: next_page link in feed
if hasattr(current_feed, ‘next_page’):
logging.info(f"Found next_page link: {current_feed.next_page}")
return current_feed.next_page
# Method 2: Atom-style pagination
if hasattr(current_feed.feed, ‘links’):
for link in current_feed.feed.links:
if link.get(‘rel’) == ‘next’:
logging.info(f"Found Atom-style next link: {link.href}")
return link.href
# Method 3: RSS 2.0 pagination (using lastBuildDate)
if hasattr(current_feed.feed, ‘lastBuildDate’):
last_date = current_feed.feed.lastBuildDate
if hasattr(current_feed.entries, ‘last’):
last_entry = current_feed.entries[-1]
if hasattr(last_entry, ‘published_parsed’):
last_entry_date = datetime(*last_entry.published_parsed[:6])
# Try to construct next page URL with date parameter
if ‘?’ in feed_url:
next_url = f"{feed_url}&before={last_entry_date.strftime(‘%Y-%m-%d’)}"
else:
next_url = f"{feed_url}?before={last_entry_date.strftime(‘%Y-%m-%d’)}"
logging.info(f"Constructed date-based next URL: {next_url}")
return next_url
# Method 4: Check for pagination in feed description
if hasattr(current_feed.feed, ‘description’):
desc = current_feed.feed.description
# Look for common pagination patterns in description
next_page_patterns = [
r’next page: (https?://\S+)’,
r’older posts: (https?://\S+)’,
r’page \d+: (https?://\S+)’
]
for pattern in next_page_patterns:
match = re.search(pattern, desc, re.IGNORECASE)
if match:
next_url = match.group(1)
logging.info(f"Found next page URL in description: {next_url}")
return next_url
return None
def get_feed_type(feed):
"""Determine if the feed is RSS 2.0 or Atom format."""
if hasattr(feed, ‘version’) and feed.version.startswith(‘rss’):
return ‘rss’
elif hasattr(feed, ‘version’) and feed.version == ‘atom10’:
return ‘atom’
# Try to detect by checking for Atom-specific elements
elif hasattr(feed.feed, ‘links’) and any(link.get(‘rel’) == ‘self’ for link in feed.feed.links):
return ‘atom’
# Default to RSS if no clear indicators
return ‘rss’
def get_entry_content(entry, feed_type):
"""Get the content of an entry based on feed type."""
if feed_type == ‘atom’:
# Atom format
if hasattr(entry, ‘content’):
return entry.content[0].value if entry.content else ”
elif hasattr(entry, ‘summary’):
return entry.summary
else:
# RSS 2.0 format
if hasattr(entry, ‘content’):
return entry.content[0].value if entry.content else ”
elif hasattr(entry, ‘description’):
return entry.description
return ”
def get_entry_date(entry, feed_type):
"""Get the publication date of an entry based on feed type."""
if feed_type == ‘atom’:
# Atom format uses updated or published
if hasattr(entry, ‘published_parsed’):
return datetime(*entry.published_parsed[:6])
elif hasattr(entry, ‘updated_parsed’):
return datetime(*entry.updated_parsed[:6])
else:
# RSS 2.0 format uses pubDate
if hasattr(entry, ‘published_parsed’):
return datetime(*entry.published_parsed[:6])
return datetime.now()
def get_feed_metadata(feed, feed_type):
"""Extract metadata from feed based on its type."""
metadata = {
‘title’: ”,
‘description’: ”,
‘language’: ‘en’,
‘author’: ‘Unknown’,
‘publisher’: ”,
‘rights’: ”,
‘updated’: ”
}
if feed_type == ‘atom’:
# Atom format metadata
metadata[‘title’] = feed.feed.get(‘title’, ”)
metadata[‘description’] = feed.feed.get(‘subtitle’, ”)
metadata[‘language’] = feed.feed.get(‘language’, ‘en’)
metadata[‘author’] = feed.feed.get(‘author’, ‘Unknown’)
metadata[‘rights’] = feed.feed.get(‘rights’, ”)
metadata[‘updated’] = feed.feed.get(‘updated’, ”)
else:
# RSS 2.0 format metadata
metadata[‘title’] = feed.feed.get(‘title’, ”)
metadata[‘description’] = feed.feed.get(‘description’, ”)
metadata[‘language’] = feed.feed.get(‘language’, ‘en’)
metadata[‘author’] = feed.feed.get(‘author’, ‘Unknown’)
metadata[‘copyright’] = feed.feed.get(‘copyright’, ”)
metadata[‘lastBuildDate’] = feed.feed.get(‘lastBuildDate’, ”)
return metadata
def create_ebook(feed_url, start_date, end_date, output_file):
"""Create an ebook from RSS feed entries within the specified date range."""
logging.info(f"Starting ebook creation from feed: {feed_url}")
logging.info(f"Date range: {start_date.strftime(‘%Y-%m-%d’)} to {end_date.strftime(‘%Y-%m-%d’)}")
# Parse the RSS feed
feed = feedparser.parse(feed_url)
if feed.bozo:
logging.error(f"Error parsing feed: {feed.bozo_exception}")
return False
# Determine feed type
feed_type = get_feed_type(feed)
logging.info(f"Detected feed type: {feed_type}")
logging.info(f"Successfully parsed feed: {feed.feed.get(‘title’, ‘Unknown Feed’)}")
# Create a new EPUB book
book = epub.EpubBook()
# Extract metadata based on feed type
metadata = get_feed_metadata(feed, feed_type)
logging.info(f"Setting metadata for ebook: {metadata[‘title’]}")
# Set basic metadata
book.set_identifier(feed_url) # Use feed URL as unique identifier
book.set_title(metadata[‘title’])
book.set_language(metadata[‘language’])
book.add_author(metadata[‘author’])
# Add additional metadata if available
if metadata[‘description’]:
book.add_metadata(‘DC’, ‘description’, metadata[‘description’])
if metadata[‘publisher’]:
book.add_metadata(‘DC’, ‘publisher’, metadata[‘publisher’])
if metadata[‘rights’]:
book.add_metadata(‘DC’, ‘rights’, metadata[‘rights’])
if metadata[‘updated’]:
book.add_metadata(‘DC’, ‘date’, metadata[‘updated’])
# Add date range to description
date_range_desc = f"Content from {start_date.strftime(‘%Y-%m-%d’)} to {end_date.strftime(‘%Y-%m-%d’)}"
book.add_metadata(‘DC’, ‘description’, f"{metadata[‘description’]}\n\n{date_range_desc}")
# Create table of contents
chapters = []
toc = []
# Process entries within date range
entries_processed = 0
entries_in_range = 0
consecutive_out_of_range = 0
current_page = 1
processed_urls = set() # Track processed URLs to avoid duplicates
logging.info("Starting to process feed entries…")
while True:
logging.info(f"Processing page {current_page} with {len(feed.entries)} entries")
# Process current batch of entries
for entry in feed.entries[entries_processed:]:
entries_processed += 1
# Skip if we’ve already processed this entry
entry_id = entry.get(‘id’, entry.get(‘link’, ”))
if entry_id in processed_urls:
logging.debug(f"Skipping duplicate entry: {entry_id}")
continue
processed_urls.add(entry_id)
# Get entry date based on feed type
entry_date = get_entry_date(entry, feed_type)
if entry_date < start_date:
consecutive_out_of_range += 1
logging.debug(f"Skipping entry from {entry_date.strftime(‘%Y-%m-%d’)} (before start date)")
continue
elif entry_date > end_date:
consecutive_out_of_range += 1
logging.debug(f"Skipping entry from {entry_date.strftime(‘%Y-%m-%d’)} (after end date)")
continue
else:
consecutive_out_of_range = 0
entries_in_range += 1
# Create chapter
title = entry.get(‘title’, ‘Untitled’)
logging.info(f"Adding chapter: {title} ({entry_date.strftime(‘%Y-%m-%d’)})")
# Get content based on feed type
content = get_entry_content(entry, feed_type)
# Clean the content
cleaned_content = clean_html(content)
# Create chapter
chapter = epub.EpubHtml(
title=title,
file_name=f’chapter_{len(chapters)}.xhtml’,
content=f'<h1>{title}</h1>{cleaned_content}’
)
# Add chapter to book
book.add_item(chapter)
chapters.append(chapter)
toc.append(epub.Link(chapter.file_name, title, chapter.id))
# If we have no entries in range or we’ve seen too many consecutive out-of-range entries, stop
if entries_in_range == 0 or consecutive_out_of_range >= 10:
if entries_in_range == 0:
logging.warning("No entries found within the specified date range")
else:
logging.info(f"Stopping after {consecutive_out_of_range} consecutive out-of-range entries")
break
# Try to get more entries if available
next_page_url = get_next_feed_page(feed, feed_url)
if next_page_url:
current_page += 1
logging.info(f"Fetching next page: {next_page_url}")
feed = feedparser.parse(next_page_url)
if not feed.entries:
logging.info("No more entries available")
break
else:
logging.info("No more pages available")
break
if entries_in_range == 0:
logging.error("No entries found within the specified date range")
return False
logging.info(f"Processed {entries_processed} total entries, {entries_in_range} within date range")
# Add table of contents
book.toc = toc
# Add navigation files
book.add_item(epub.EpubNcx())
book.add_item(epub.EpubNav())
# Define CSS style
style = ”’
@namespace epub "http://www.idpf.org/2007/ops";
body {
font-family: Cambria, Liberation Serif, serif;
}
h1 {
text-align: left;
text-transform: uppercase;
font-weight: 200;
}
”’
# Add CSS file
nav_css = epub.EpubItem(
uid="style_nav",
file_name="style/nav.css",
media_type="text/css",
content=style
)
book.add_item(nav_css)
# Create spine
book.spine = [‘nav’] + chapters
# Write the EPUB file
logging.info(f"Writing EPUB file: {output_file}")
epub.write_epub(output_file, book, {})
logging.info("EPUB file created successfully")
return True
def main():
parser = argparse.ArgumentParser(description=’Convert RSS feed to EPUB ebook’)
parser.add_argument(‘feed_url’, help=’URL of the RSS feed’)
parser.add_argument(‘–start-date’, help=’Start date (YYYY-MM-DD)’,
default=(datetime.now() – timedelta(days=365)).strftime(‘%Y-%m-%d’))
parser.add_argument(‘–end-date’, help=’End date (YYYY-MM-DD)’,
default=datetime.now().strftime(‘%Y-%m-%d’))
parser.add_argument(‘–output’, help=’Output EPUB file name’,
default=’rss_feed.epub’)
parser.add_argument(‘–debug’, action=’store_true’, help=’Enable debug logging’)
args = parser.parse_args()
if args.debug:
logging.getLogger().setLevel(logging.DEBUG)
# Parse dates
start_date = datetime.strptime(args.start_date, ‘%Y-%m-%d’)
end_date = datetime.strptime(args.end_date, ‘%Y-%m-%d’)
# Create ebook
if create_ebook(args.feed_url, start_date, end_date, args.output):
logging.info(f"Successfully created ebook: {args.output}")
else:
logging.error("Failed to create ebook")
if __name__ == ‘__main__’:
main()
[/python]